top of page

The AI Ethics Class - Musings on Potential Harms

  • O Leonard
  • Feb 9, 2025
  • 2 min read

This information originally was authored in February 2025 as part of an online discussion. It is has been edited for clarity and formatting.


Microsoft is the leading provider of productivity software for enterprise and government, so if and when there is a breach, there is the potential for large amounts of data to be leaked.


Recently, Forbes reported that the Government Efficiency team is using Azure AI services, the technology behind Copilot, to analyze Treasury data. So, in addition to the fallout among public sector workers and vendors, there is potential to affect anyone who receives Treasury payments.


A robot hand reaching out
Robot Hand Reaching Out

Although the use of AI may identify hidden relationships and improve system efficiency, if this information is leaked and combined with other data (e.g. OPM data), sensitive data will be compromised, directly or through inference.  Once released, even after the original breach is repaired, the information will be available on the internet forever.


The Next Breach


There have already been several breaches, two notable ones involving OPM data occurring in 2021 and 2024.  However, these occurred before the Copilot historical searches were stored. Any future breach, could affect most US citizens and Western nations who use Microsoft 365.

It appears that very little can be done to prevent this outcome. The more secure a product is, the less user friendly it becomes. As mentioned in a Congressional hearing on regulating AI, the technology is extremely expensive and tech companies are therefore incentivized to make it as user friendly as possible to encourage adoption and a means to recover their large investments in it.


Prior to OpenAI releasing ChatGPT to the public, companies had to apply with Microsoft to use their most advanced AI products and show they would use the technology responsibly. Now, this (ChatGPT) and other generative AI systems are readily available to the public without any gatekeeping.


Is Synthetic Data the Answer?


To mitigate some of these harms, some have championed the use of synthetic data. The use of synthetic data may address some of the risks associated with re-identification and the lack of new data, but it creates new risks for model drift and compounding biases. Although human sentiment data would now be available to larger audiences, it creates a scenario where, depending on the cost, smaller, mid-size, and publicly traded companies feel pressure to use synthetic data to realize cost savings.


Any new data would be limited to topics of interest to the largest companies or private investors. Some scholars have pointed out that there is a potential for AI generated date to become a commodity and "authentic" data to become a luxury in much the same way that handmade products became reserved for the wealthy after the industrial revolution. As the market for this type of data shrinks, fewer people will have the expertise to gather it and to identify issues with synthetic data. Returning to the example of clothes: After the rise of fast fashion, although they still exist, it is now more difficult for the average person to find a seamstress or shoe repair shop.

Comments


bottom of page