Privacy and Data Agency AI Benchmarking Scale

Date: August 11, 2026

Read the full PDF here: Privacy and Data Agency AI Benchmarking Scale

Brief

As artificial intelligence (AI) models become more advanced and more widely deployed, data collection and storage becomes a more pressing ethical and regulatory issue. Increased data collection, for AI model training, operation, or general advertising, has resulted in more frequent and severe violations of privacy [1]. The Privacy and Data Agency Benchmarking Scale was created as a way to standardize privacy-enhancing practices across the technological industry, particularly regarding AI development and deployment. This scale outlines the privacy practices data collectors should follow in order to preserve the privacy of users, individuals who interact with their products.

This scale is designed to assess and compare AI models based off of how well they adhere to privacy practices. Using this scale, AI users can weigh models against each other, having the choice to understand and select which AI model may align most with their values and interests. This scale aims to create a “marketplace” where companies compete to win consumer favor through having a model that is most ethically aligned with their target users. Furthermore, this scale has far-reaching applications particularly in the legal field; it can be used by lawmakers and regulatory bodies as a baseline framework upon which privacy legislation can be built.

Our Mission

Guiding the development of AI tools to uphold democratic values and champion human rights.

We believe that democracy is more than a political system; it is a dynamic modus vivendi that extends beyond politics and permeates all aspects of society. Just as the US Constitution is regarded as a living document, democracy is a living system that engages a variety of stakeholders and enables individuals to take part in determining the future of our world. In short, democracy allows all of us to work together to build the future.

We believe AI needs to be created under a democratic framework. Firstly, the AI industry must be governed democratically, and secondly, AI must promote and uphold to democratic values. AI technologies that are governed democratically must have mechanisms that allow the public to be properly represented and they must have systems of transparency and accountability. AI that supports democratic values helps users maintain autonomy and self-determination, helps promote financial opportunity and equity, and encourages the free flow of information. To further understand the aspects of Democratic AI, see a full list of our Democratic AI Principles here.

The Privacy and Data Agency AI Benchmarking Scale contributes to the development of democratic AI by providing informational transparency, fostering individual autonomy, and establishing a mechanism through which the broader public can influence the development of AI tools. More specifically, this scale assists AI users in understanding the privacy protection levels of different AI models, and allows users to choose how they engage with AI models.

In addition to simply providing more information about AI models and increasing AI technical transparency, this scale will foster the development of a competitive “market” of AI tools that users can “shop”. Having AI developers report their data privacy ratings will encourage them to compete on standards of privacy, similarly to the ways that car companies compete on miles per gallon (mpg) ratings. Users will be able to utilize the power of their consent, and can decide with their actions which AI models they wish to engage with depending on their privacy protection levels. For example, an AI model with a rating of “A” has more robust data privacy protection practices than an AI model with a rating of “C”; if all else is equal amongst the models, if the models have similar capabilities, users who value privacy may opt to utilize the model with a rating of “A” over “C”. AI developers will therefore be incentivized to compete on standards of privacy in order to obtain and retain users.

This scale will also provide policy makers and regulators with information needed to track how well AI developers are creating technologies in accordance with legal guidelines. While there does not yet exist any nationwide regulation regarding AI, data, and privacy, multiple states have begun to introduce state-wide measures to enhance user privacy or encourage corporate transparency regarding AI model training data. The legislators in these states can utilize this benchmarking scale as a means to track how well AI developers are creating technologies in accordance with local guidelines.

Democracy relies on an informed and engaged populace; this scale informs users on the state of data privacy protections in AI models and increases their access to information regarding AI, data, and privacy. This scale enables the public to understand the impact of the AI models around them; an informed public can make informed decisions and take ownership of the future of technology.

The Foundation

Data privacy is a critical concern in AI development; AI tools consume unprecedented amounts of data for their development and deployment. However, gaining access to this enormous amount of data has put privacy considerations at risk. IBM's AI research division used approximately 1 million public Flickr photos to train facial recognition models without photographer or subject consent [2]. Clearview AI scraped over 30 billion images without consent to build its facial recognition database, drawing millions in fines. Clearview AI reported being used by American police over one million times [3].

These breaches of privacy directly concern a large percentage of Americans, who are worried about AI’s detrimental effect on data privacy. Approximately 61% of Americans consider it important to limit who has access to their data [4]. Around 80% are concerned that their personal data is being used to train AI models [5]. Americans translate their concerns into action; 85% of American consumers have actively taken steps to protect themselves from such data leaks, breaches of data, or consensual use of their data to train algorithms [6]. However, privacy fatigue is real and has an effect on how people view their data and their own measures to try and protect it [7]. It is crucial that leaders enact measures to protect data privacy in the age of AI and institute mechanisms that bolster users’ abilities to control their own data.

In the current digital environment privacy boundaries are frequently violated. Individuals often agree to privacy policies they do not understand. An individual’s data is collected even after opting out of data sharing, either through data inference, profiling, or simply having their opt-out preferences ignored. It is hard, if not impossible, to truly avoid data collection.

This creates an atmosphere of data exploitation, an atmosphere where consent cannot thrive. In spaces where individuals do not have control over their data there exists an imbalance of power, access, and control over information.

It is therefore beneficial to create a benchmark to judge an AI model’s adherence to data privacy practices, data practices that promote a flourishing data environment. This benchmark differs from a law, such as the Children's Online Privacy Protection Act (COPPA). While laws outline boundaries and are accompanied by enforcement mechanisms, this scale instead establishes a standard of data privacy that institutions should aspire to meet. This benchmark outlines what comprehensive data privacy protection measures look like. Using this scale:

(i) Consumers gain an understanding of what comprehensive data privacy protection measures look like.

(ii) Lawmakers have a baseline to work from when drafting data privacy and protection bills.

(iii) Institutions have a detailed list of measures they can adopt in order to increase their data privacy protection practices.

There exist a multitude of benchmarks that measure the technical capabilities of AI models, as well as scales that assess how well AI companies align with Responsible AI practices. Existing benchmarks include MMLU, GPQA, AIME, and SWE-bench Verified that gauge model capability. These benchmarks test a model’s aptitude for math, software engineering tasks, logical reasoning, and web tasks, amongst other capabilities [8]. Benchmarks that assess Responsible AI alignment include BBQ (2021), measuring fairness and bias; INTIMA for companionship tracking; KaBLE for belief vs fact tracking; HarmBench (2024), Cybench (2024), StrongREJECT (2024), and WMDP (2024), measuring security; HHEM and SimpleQA (2024), measuring factuality and truthfulness; and MakeMePay (2024), measuring autonomy and human agency [9]. Privacybench measures the release of secrets [10].

However, there does not yet exist a benchmark or scale that measures how well an AI model respects an individual’s right to privacy, and how ethically the AI system handles data. This framework attempts to fill this gap by creating a scale by which one can assess how well AI systems uphold and adhere to principles of data privacy and practices that foster individual control over data. This scale is part of a larger body of work by the Center for Democratic AI to create an AI Democracy Benchmark that evaluates how well AI models align with democratic principles and promote democratic values.

The goal of this framework is threefold:

  1. Empower user choice; an informed public can make informed decisions regarding their use of AI models. Users can influence, with their decisions on what AI models they choose to engage with, the future of the AI industry.

  2. Foster ethical AI model creation by incentivizing AI developers to employ strong data privacy protection practices.

  3. Establish a standard set of data privacy protection practices that can be implemented across institutions, particularly in the AI industry. With the implementation of this scale, the concept of data privacy graduates from an abstract notion of best practices to a measurable, traceable set of criteria. Regulators, corporations, and users alike will have the ability to track the progress of data privacy protections within the AI industry over time, assess the practices of individual AI models, and conduct analyses on the impacts of data privacy protections.

The Framework

Section 1. Scale Metrics

For the purposes of this benchmark, the following definitions shall apply:

a. “Privacy” means protection of individuals’ confidentiality, anonymity, informed consent, and control over personal data across the AI life cycle (collection, training, deployment, reuse).

b. “Data collectors” refers to any entity that collects information. This includes but is by no means limited to developers of AI models, the host of a website, the company running a social media platform, manufacturers of technological devices such as smartphones or doorbells, owners of CCTV cameras, and retail stores.

c. “User” refers to any person who interacts with a data collector.

d. “Data” refers to any information, including both digital and offline information. This includes but is not limited to likeness, digital metadata, location, behavioral patterns, health data, images, demographics, and identifying information.

This framework evaluates how well an AI model protects the privacy of users in the handling of their data. The data collectors who create such a model should make all reasonable attempts to adhere to the following stipulations in order to most comprehensively promote principles of privacy:

Active opt in. Data collectors must by default collect no data from a user, placing the user in a default “opt out of data collection” status, thereby allowing users to opt in rather than out of collection. By having opt out be the default data sharing status, data collectors require users to make an active choice to opt in, strengthening active consent.

Local data storage. Data collectors must be encouraged to refrain from automatically collecting consensually-shared user information into their own databases. Rather, collectors should be encouraged to allow user information gathered from a device to stay on the local device, referencing that data only when need be. This reduces reliance on data centers, decentralizes data storage, and increases privacy by only referencing data when needed, as opposed to collecting it for unnecessary and/or secondary purposes.

Security Safeguards. Personal data should be protected by reasonable security safeguards against risks such as loss or unauthorized access, destruction, use, modification or disclosure of data. This increases safety around privacy measures, reducing the risk of data breaches and user data being shared beyond their consent.

Open Access. Users should be able to easily access the data collected on themselves. Means should be readily available for establishing the existence and nature of personal data, and users should be able to easily see the existence and nature of their data. This increases transparency and user understanding and control over their shared data.

User Communication. Users should have the right to obtain from a data collector confirmation of whether or not the data collector has data relating to them, and request access to such data. Confirmations must arrive within a reasonable time in a form that is readily intelligible to them. If that request is denied, they must be able to challenge such denial. This increases accountability, clarity, communication, as well as user control over their privacy preferences.

Transparency. Data collectors must explain how and why data is used and who else it is shared with. This enables a user to make an informed choice about how their data is utilized, thereby giving them more control, information, and ability to actively consent.

Equal treatment. Data collectors must provide equal services and treatment to users who do and do not opt in to data sharing. This ensures that users are not penalized for opting out of sharing their data, making certain that they are not consenting to their data usage under pressure or lack of viable alternative. Policies that mandate unnecessary sharing of user data in order to access a product or service are not in alignment with principles of privacy and consent. Users should not be denied access to a product or service because they did not share their data. This applies to data beyond what is reasonably necessary for the product or service to function.

Partner transparency. Data collectors must explain who they share collected data with and the purposes for which the data is shared. Users must remain informed of and understand who their data is shared with and why it is shared in order to consent to this sharing. Companies must make all reasonable effort to disclose their data partners and provide options for users to refuse without penalty for their data to be shared with them.

Equal reconsent. Data collectors must ask for updated consent from those who opt in to data sharing and those who opt out of data sharing at equal frequencies. Consent expires and must be reaffirmed after a period of time. Data collectors must reaffirm a user's data sharing consent or lack thereof. In addition, they must do so equally, no matter whether a user has previously decided to share or not share their data. This will prevent undue pressure and excess fatigue on users who have opted out of data sharing. Data collectors who repeatedly ask those who have opted out of data sharing to reaffirm their preferences may fatigue them into opting in, especially if these data sharing requests occur at a higher frequency than the requests for data preferences of those who had already opted in.

Active reconsent. For each privacy policy update, users must actively reconsent to new data usages. Data collectors must not implement new privacy conditions until the user has actively affirmed the new conditions. This places agency with the user, and ensures that no changes to previously agreed-upon data uses and privacy agreements are put into effect without the user’s explicit understanding and agreement.

Plain English policies. Data collectors must provide a clear, understandable, brief plain English summary of data sharing policies and updates for users to consent to. In addition to a longer and more detailed privacy policy statement, data collectors should have a shorter, understandable, abridged version that all users are able to read and comprehend. This must contain the most important details of data sharing, why data is used, who it is shared with, and how one can remove their data from any applied practice. This allows users to easily gain a deeper understanding of the data uses they are consenting to, and prevents deception or misunderstandings that may occur when users are presented only with long, nuanced policy privacy policies that are difficult to understand.

AI Training Transparency. Data collectors must make users aware that their data is being used to train AI models or other AI training or testing purposes, if applicable. Any AI model trainer must ensure that the dataset they use is made of data sources who have given their consent to have their data used for this model training. The onus is on developers to ensure that the individuals in the dataset they use for training are updated with the new use of their data. This combats the idea of one-time consent, where an individual gives consent for a company to use their data and the company can use if for multiple purposes without updating the user. The user must be made aware of new uses for their data, e.g. training a new AI model, so that they can consent to that new use. If a developer sources their dataset from a different entity, those two entities must make the users in the dataset aware of the use for AI training.

Necessary data collection only. Data collectors must collect only the data necessary for specific operations, and should not perform extraneous operations that are not central to the collector’s mission or purpose. Data collectors must also explain how and why the data they are collecting is necessary for a specific operation. This maximizes privacy by reducing data shared as much as possible.

Collection-free usage. Users should be able to utilize services offered by data collectors with the option to not have their usage data to be used to improve the data collector’s services. This increases privacy by allowing people to use AI or digital services without being continuously monitored.

Synthetic data transparency. AI developers should report the quantity and nature of synthetic data used in their model’s training and operations, as well has how the data was synthesized. This increases user understanding of model creation and data transparency.

Anti-Coercion Safeguards. Data collectors must ensure that no coercive measures are taken to influence users to share their data unwillingly. Coercive measures include but are not limited to: pre-selected opt-in to data sharing boxes, hidden opt-outs, visually biased opt-in and opt-out menus, and service blocking. Consent cannot be given under coercive conditions; preventing coercion allows for true consent to be given and ensures that privacy is determined by a user's wishes.

Ease of data removal. Data collectors must enable users to permanently have their data removed and forgotten easily and in a timely manner. This makes revoking previously shared consent impactful, allowing people to obtain privacy at any time.

Equal ease of consenting, denying, and revoking consent. Withdrawal from data sharing or declining to share data must be as easy as granting access to data. If a user is able to consent to sharing data with one click, declining to share data or revoking that consent must also be able to be done with one click. This brings down inhibitions or barriers to users being able to opt out of data sharing and take back their consent.

Proof of auditing. Data collectors must undergo regular auditing of its practices by an independent board. Data collectors must be prepared to submit their privacy management infrastructure for an audit at any time. This ensures that practices are being followed, it increases accountability, transparency, and enforcement of data privacy practices.

Data consent management infrastructure. Data collectors must have consent management infrastructure that tracks not just the current state (for example, yes or no) of agreement to data collection and sharing, but the history of consent as well: when consent was given, for which specific purposes, under which version of a privacy policy, and when it was withdrawn or expired. This produces an auditable trail of information so that independent regulators can verify and enforce proper data protection practices.

Reverse-engineering protection measures. Data collectors must not reverse engineer or infer data and information that a user has declined to share. This stipulation goes beyond enforcing the law of data privacy and attempts to enforce the spirit of data privacy. Even if a company does not collect data from a user, they must not take extraneous measures to obtain said data or generate substitute data from means such as inferencing.

Section 2. Scale Application

The scale has 21 metrics that are assessed on a “Completed/Not Completed” basis. Each metric is assigned one point, so the maximum number of points an entity can obtain on this scale is 21 points. For ease of use and understanding, however, the scale will be converted to a letter-grading system, with point ranges determining an entity’s letter grade. The letter grade will be as follows:

A+: 19-21 points

A: 16-18 points

B+: 13-15 points

B: 10-12 points

C: 7-9 points

D: 3-6 points

F: 0-2 points

Seeing as there is limited oversight and regulatory boards to independently assess the practices of an AI company, grades will be determined on a self-reporting basis. Companies will report their data privacy practices, which will be used to determine their grade.

Conclusion

It is crucial that AI policy makers, developers, and users have a standard through which they can assess how well AI models promote principles of privacy. The preceding framework is our first contribution to that effort, but only the beginning.

Progress in promoting privacy in technology will depend on adopting this scale widely and increasing transparency across the AI industry. Data collectors, especially AI developers, are at present not forthcoming nor transparent about their Responsible AI metrics. This must change if principles of democracy are to be upheld in the AI industry; AI developers must enable regulators, auditors, and users to have access to the information about their products and systems.

A standardized system of oversight should be developed; there should exist an independent oversight board that evaluates and audits companies adherence to privacy measures and privacy practices. Policymakers can work on the creation of laws and sanctions that complement these principles, boundaries, and actions.

Increased transparency and information sharing around data privacy practices will enable this scale to be more concretely attuned to the state of AI and data collection. There needs to be constant monitoring of the state of privacy practices and the effect of data collection on the stability of democracy.

This scale will help promote democracy. It will help protect the freedom to privacy that every individual is owed; increase the public’s understanding of AI models and their impact on data; give users who interact with emerging technology more of a say in how their data is used, and what values should drive the future of technological innovation. This scale puts more power into the hands of the public.

References

  1. See Csupo v. Alphabet Inc., No. 19CV352557 (Cal. Super. Ct. Santa Clara Cnty., jury verdict July 1, 2025); Clearview AI, Inc., Consumer Privacy Litigation, No. 1:21-cv-00135, MDL No. 2967 (N.D. Ill., final approval Mar. 20, 2025); Facebook Biometric Information Privacy Litigation, No. 3:15-cv-03747 (N.D. Cal., final approval Feb. 26, 2021); United States v. Amazon.com, Inc., No. 2:23-cv-00811 (W.D. Wash., filed May 31, 2023); Dinerstein v. Google, LLC, 73 F.4th 502 (7th Cir. 2023); Vance v. International Business Machines Corp., No. 1:20-cv-00577 (N.D. Ill.).

  2. "IBM Used Flickr Photos for Facial-Recognition Project," BBC News, March 13, 2019.

  3. James Clayton and Ben Derico, "Clearview AI Used Nearly 1m Times by US Police, It Tells the BBC," BBC News, March 27, 2023.

  4. Shinmin Bali, "Data Privacy Day US 2026: How Concerned Are Americans about Data Security?," YouGov, January 13, 2026.

  5. John Koetsier, "Americans Are Terrified about AI: 80% Say AI Will Help Criminals Scam Them," Forbes, August 22, 2023.

  6. "New Deloitte Survey: Increasing Consumer Privacy and Security Concerns in the Generative AI Era," Deloitte, December 2, 2024.

  7. Hanbyul Choi, Jonghwa Park, and Yoonhyuk Jung, "The Role of Privacy Fatigue in Online Privacy Behavior," Computers in Human Behavior 81 (2018): 42–51.

  8. "Benchmarks," Epoch AI, accessed August 11, 2026.

  9. Stanford Institute for Human-Centered Artificial Intelligence, "Responsible AI," chap. 3 in The 2026 AI Index Report (Stanford, CA: Stanford University, 2026).

  10. Srija Mukhopadhyay, Sathwik Reddy, Shruthi Muthukumar, Jisun An, and Ponnurangam Kumaraguru, "PrivacyBench: A Conversational Benchmark for Evaluating Privacy in Personalized AI," arXiv, December 31, 2025.

General References:

DataGrail, "What Is Opt-Out and Opt-In Consent?," DataGrail (blog), February 26, 2026.

"Data Protection in the AI Era: Benchmarking EU GDPR and AIA, to Reform Saudi Data Protection Law," Cogent Social Sciences, April 7, 2026.

Previous
Previous

Data Breaches and Privacy Violations: A Summary