T-03
Transparency & Disclosure
Training Data Disclosure
Developers must disclose information about the data used to train AI models. Public disclosure obligations require posting documentation on the developer's website covering dataset sources, data types, volume, IP status, personal information presence, processing methods, collection timeframes, and use of synthetic data. Regulator disclosure obligations require submitting similar documentation to a designated authority, which may treat it as confidential.
Sub-obligations3
Bills40
Jurisdictions17
Enacted2
Show
Sort bills within section

3 sub-obligations of T-03

Click any row to jump to its bills below.
ID Sub-Obligation Enacted Live Failed Total
T-03.1 Regulator disclosure
Developers must provide training data documentation to designated regulatory authorities. This disclosure may be treated as confidential and is not required to be made public.
0Enacted 1Live 3Failed 4Total Jump →
T-03.2 Public disclosure
Developers must post training data documentation publicly on their website before making a system available and before each new release or substantial modification. Substantial modification includes retraining and fine-tuning.
1Enacted 13Live 2Failed 16Total Jump →
T-03.3 Training Data Governance Disclosure to Deployers
Developers must disclose to deployers the data governance measures applied to training datasets, including examination of data source suitability, possible biases, and mitigation steps taken, as part of pre-deployment technical documentation.
1Enacted 9Live 9Failed 19Total Jump →
Bills That Map This Requirement 53 mappings
T-03.1
Regulator disclosure
Developers must provide training data documentation to designated regulatory authorities. This disclosure may be treated as confidential and is not required to be made public.
Enacted
0
Live
1
Failed
3
Total
4
US
Introduced
Any person who uses a training dataset containing copyrighted works to train or release a generative AI model must submit a notice to the Register of Copyrights containing (1) a sufficiently detailed summary of each copyrighted work in the dataset and (2) the URL for the training dataset if publicly available. For models first used or released after the effective date, notice must be filed at least 30 days before commercial use or release; for pre-existing models, notice is due within 30 days after the Register issues implementing regulations.
TX
TX HB 1709 (AI Governance) § Bus. & Com. Code § 551.003
Failed
Developers must maintain detailed records of all generative AI training data consistent with the NIST AI RMF GenAI Profile (GV-1.2-007).
US
Failed
Covered entities must maintain and keep updated documentation of all data or input information used to develop, test, maintain, or update the automated decision system, including data sourcing, methodology, consent practices, representativeness, and quality measures.
US
Failed
Any person who creates or significantly alters a training dataset used in building a generative AI system must submit a notice to the Register of Copyrights containing a sufficiently detailed summary of copyrighted works used and the dataset URL (if publicly available), filed at least 30 days before the system is made available to consumers (for new systems) or within 30 days of the Act's effective date (for existing systems).
T-03.2
Public disclosure
Developers must post training data documentation publicly on their website before making a system available and before each new release or substantial modification. Substantial modification includes retraining and fine-tuning.
Enacted
1
Live
13
Failed
2
Total
16
CA
Enacted eff 2026-01-01
Developers must post on their website, by January 1, 2026, and before each subsequent public release or substantial modification, documentation of the training data used for any generative AI system or service released on or after January 1, 2022, including dataset sources, data types, volume, IP status, personal information presence, processing methods, collection timeframes, and use of synthetic data. Exemptions apply for systems whose sole purpose is security and integrity, national airspace operations, or national security/military/defense systems available only to federal entities.
NY
Engrossed
Developers must post on their website, on or before January 1, 2027, and before each subsequent public release or substantial modification of a generative AI model or service released on or after January 1, 2022, documentation regarding the data used to train the model or service, including a high-level summary of the datasets covering: (1) the sources or owners of the datasets; (2) a description of how the datasets further the intended purpose; (3) the number of data points, which may be in general ranges with estimates for dynamic datasets; (4) a description of the types of data points (for labeled datasets, the types of labels used; for unlabeled datasets, general characteristics); (5) whether the datasets include data protected by copyright, trademark, or patent or are entirely in the public domain; (6) whether the datasets were purchased or licensed; (7) whether the datasets include personal information or personal identifying information; (8) whether the datasets include aggregate consumer information; (9) whether there was any cleaning, processing, or other modification, including its intended purpose; (10) the time period during which data was collected, including notice if collection is ongoing; (11) the dates datasets were first used during development; and (12) whether the model uses or continuously uses synthetic data generation. This obligation does not apply to models solely for national airspace aircraft operations or models developed for national security, military, or defense purposes available only to a federal entity.
GA
GA HB 1603 (AI Performer Protection) § O.C.G.A. § 10-1-972
Introduced eff 2027-01-01
Production companies deploying AI systems for use in production in Georgia must issue a disclaimer describing how AI was adopted and deployed and identifying any data, sources, or metrics that were used.
NY
Introduced
Developers must post on their website, before January 1, 2027 and before each subsequent public release or substantial modification of a generative AI system released on or after January 1, 2022, detailed information about journalism content used in training: (1) URLs accessed by crawlers, (2) a description of the content used sufficient to identify individual works, including type, provenance, and means of acquisition, (3) whether source identifiers, terms, or copyright notices were removed, and (4) the timeframe of data collection. This obligation is excused where an express written agreement authorizes the developer's access and both parties agree not to post the information.
NY
Introduced
Developers who deploy crawlers must publicly disclose, in a manner clearly accessible to website operators and updated in real time with any changes, the identity and technical details of each crawler: (1) crawler name, IP address, and user-agent identifier, (2) the legal entity responsible, (3) the specific purposes for which each crawler is used, (4) legal entities receiving scraped data, and (5) a single point of contact for third-party communications and complaints.
NY
Introduced
Developers must post on their website, on or before January 1, 2027, and before each subsequent public release or substantial modification of a generative AI model or service released on or after January 1, 2022 and made available to New Yorkers, documentation regarding the data used to train the model or service. The documentation must include a high-level summary of the training datasets covering at least twelve categories: (1) sources or owners of the datasets; (2) how the datasets further the model's intended purpose; (3) the number of data points (in general ranges, with estimates for dynamic datasets); (4) the types of data points (label types for labeled datasets, general characteristics for unlabeled datasets); (5) whether the datasets include data protected by copyright, trademark, or patent, or are entirely in the public domain; (6) whether the datasets were purchased or licensed; (7) whether the datasets include personal information or personal identifying information; (8) whether the datasets include aggregate consumer information; (9) whether there was any cleaning, processing, or other modification, and the purpose of those efforts; (10) the time period during which data was collected, including notice if collection is ongoing; (11) the dates datasets were first used in development; and (12) whether the model uses or continuously uses synthetic data generation. Exemptions apply for generative AI models whose sole purpose is aircraft operation in the national airspace, and for models developed for national security, military, or defense purposes available only to a federal entity.
NY
Introduced
Developers must post on their website, by January 1, 2027 and before each subsequent public release or substantial modification, detailed information about journalism content used for AI utilization — including URLs accessed, content descriptions sufficient to identify individual works, whether source identifiers or copyright notices were removed, and data collection timeframes. This obligation does not apply where the developer has an express written agreement with the journalism provider authorizing content access and both parties agree not to post the information.
NY
Introduced
Developers deploying crawlers must, by January 1, 2027, publicly disclose on an easily accessible platform the identity of each crawler (name, IP address, user-agent identifiers), the responsible legal entity, specific purposes, downstream data recipients, and a single point of contact for complaints. This information must be kept current.
US
Introduced
Covered entities must disclose training data sources, documentation, testing methodology, data collection during inference, and operational information for each foundation model, both before commercial deployment and throughout the system lifecycle.
US
Introduced
Covered entities must disclose a sufficiently detailed summary of training data sources, collection methods, inference-time data collection and retention practices, and the size and composition of training data including demographic, language, and attribute information, while accounting for privacy.
US
Introduced
Covered entities whose foundation model is derived from or built upon another covered entity's foundation model must independently comply with all transparency regulations for any significant change, retraining, or adaptation they apply to the base model.
VA
Introduced
Developers must post on their website, within 72 hours of making a generative AI system available in Virginia and within 72 hours after each significant update, detailed documentation for each training dataset covering: dataset name, source, volume, IP status, Do Not Train data handling, personal data management, illegal material screening, collection timeframe, and synthetic data generation use. Systems whose sole purpose is security or integrity are exempt.
WA
Introduced
Developers must post on their website, by January 1, 2026 and before each subsequent public release or substantial modification, documentation regarding the data used to train any generative AI system or service released on or after January 1, 2022 and made publicly available to Washington residents. The documentation must include a high-level summary of training datasets covering: (1) the sources or owners of the datasets; (2) how the datasets further the system's intended purpose; (3) the number of data points (general ranges and estimates permitted for dynamic datasets); (4) the types of data points within the datasets; (5) whether the datasets were purchased, licensed, or publicly available; (6) whether the datasets include personal information as defined in RCW 19.373.010; (7) whether the datasets include aggregate consumer information; (8) whether any cleaning, processing, or modification was performed and its intended purpose; (9) the dates the datasets were first trained or last significantly updated; and (10) whether the system used or continuously uses synthetic data generation. This obligation does not apply to systems whose sole purpose is security and integrity, aircraft operation in the national airspace, or national security/military/defense purposes available only to a federal entity.
WA
Introduced eff 2027-01-01
Developers must post on their website, by January 1, 2027 and before each subsequent public release or substantial modification, documentation of the datasets used to train each generative AI system released on or after January 1, 2022 — including dataset sources, data types, volume, IP status, personal-information presence, acquisition method, processing steps, training dates, synthetic-data use, and CSAM removal steps. Exempt: security-only systems, aviation systems, national-security/defense systems available only to federal entities, and FDA-regulated systems.
MD
MD HB 823 (GenAI Training Data Transparency) § Md. Code, State Fin. & Proc. § 3.5-807
Failed
Developers must publish on their website, by January 1, 2026, and before each subsequent release or substantial modification, documentation detailing the data and datasets used to train the generative AI system — including data sources, purpose alignment, dataset size, label types, IP status, licensing status, presence of personal information and aggregate consumer information, processing methods, collection timeframes, first use during development, and synthetic data generation.
US
Failed
Covered entities must disclose training data sources (including personal data collection and information to assist copyright owners), data size and composition (including demographics and language information), data governance procedures (including editing and filtering), and data labeling methods and validation, as specified in FTC regulations.
T-03.3
Training Data Governance Disclosure to Deployers
Developers must disclose to deployers the data governance measures applied to training datasets, including examination of data source suitability, possible biases, and mitigation steps taken, as part of pre-deployment technical documentation.
Enacted
1
Live
9
Failed
9
Total
19
CO
Enacted eff 2026-05-14
Developers must, on and after January 1, 2027, make available to each deployer of a covered ADMT technical documentation that is reasonably understandable and that protects trade secrets, including: (1) a general statement of intended uses and known harmful or inappropriate uses; (2) a description of the categories of data, including personal data, used to train the ADMT (to the extent known); (3) known limitations, risks, and circumstances in which the ADMT should not be used; (4) instructions for the deployer's appropriate use, monitoring, and meaningful human review; and (5) information reasonably necessary for the deployer to comply with its disclosure obligations under § 6-1-1704. If any information is withheld, the developer must notify the deployer.
GA
Introduced
Developers must make available to deployers and other developers, to the extent feasible, all information required to be submitted to the Attorney General, plus documentation through artifacts such as model cards, data set cards, or other impact assessments necessary for the deployer or a contracted third party to complete an impact assessment. A developer that also serves as deployer is exempt from this subsection unless the system is provided to an unaffiliated deployer.
IA
Introduced
Developers must make available to deployers documentation covering: foreseeable and harmful uses, training data summaries, known limitations and discrimination risks, purpose and intended uses, intended outputs, intended benefits, bias evaluation and mitigation measures, recommended use and monitoring guidance, and deployer monitoring instructions.
MA
Introduced
Developers must provide deployers with documentation including: (1) a summary of intended and foreseeable uses of the AI system, (2) known limitations and risks including algorithmic discrimination, and (3) information on the datasets used for training including measures taken to mitigate biases.
MO
Introduced
Vendors must disclose to the contracting entity all data elements collected, all third-party recipients, all embedded libraries, software development kits, and analytics tools, all device-level permissions, and all AI components and functions.
NJ
Introduced
Vendors must provide the auditor or Department with full documentation of the AEDS design, development, training-data sources, intended purpose, outputs, economic and employment-reduction impacts, and accuracy, reliability, validity, and error-rate analyses needed to conduct the impact assessment.
NJ
Introduced
Vendors must provide the Department full documentation of the ABSDS design, development, training-data sources, intended purpose, outputs, economic and employment-reduction impacts, and accuracy, reliability, validity, and error-rate analyses needed to conduct the impact assessment.
NY
Introduced
Developers must make available to each deployer or downstream developer the following documentation for each high-risk AI decision system: (1) a general statement describing reasonably foreseeable uses and known harmful or inappropriate uses; (2) documentation disclosing training data type summaries, known limitations including algorithmic discrimination risks, system purpose, intended benefits and uses, and any other information necessary for the deployer's compliance; (3) documentation describing pre-market evaluation methods for performance and discrimination mitigation, data governance measures covering training datasets, intended outputs, discrimination-mitigation measures, and instructions for use, non-use, and human monitoring during consequential decisions; and (4) any additional documentation reasonably necessary to help the deployer understand outputs and monitor for discrimination risk. Trade secrets and security-sensitive information are exempt from disclosure.
US
Introduced
Covered entities must maintain and keep updated documentation of all data and input information used to develop, test, maintain, or update the covered algorithm, including data sourcing and licensing details, collection methodology, consumer consent practices, rationale for data use, dataset representativeness, and data quality measures.
US
Introduced
Covered entities must disclose a description of their data governance procedures for their foundation models.
CO
Failed
Developers must make available to deployers or other developers of their high-risk AI systems the documentation required under the statute, including information about the system's intended uses, known limitations, and discrimination risks.
CO
Failed
Developers must provide deployers, to the extent feasible, the documentation and information — through model cards, dataset cards, or impact assessments — necessary for deployers to complete their own impact assessments.
CO
Failed
Developers must make available to deployers or other developers the documentation specified under § 6-1-1702(2), effective June 30, 2026.
CT
Failed
Developers of general-purpose AI models must create, maintain, and make available to downstream integrators documentation enabling them to understand the model's capabilities and limitations and comply with the act, including technical integration requirements, training data descriptions, and model metadata.
IL
Failed
Developers must provide deployers with a statement of intended uses and documentation covering (1) known limitations including foreseeable algorithmic discrimination risks, (2) the types of data used to program or train the tool, and (3) how the tool was evaluated for validity and explainability before sale or licensing.
NM
Failed
Developers must make available to recipients the information necessary for impact assessments, including model cards, dataset cards, and previous impact assessments relevant to the system.
US
Failed
Developers of high-impact AI systems must provide deployers with the information reasonably necessary for the deployer to comply with transparency reporting requirements, including an overview of training data (size, sources, copyrighted data, PII), documentation of the baseline model's structure and context (input/output modality, model size, architecture), known capabilities, limitations, and risks, and downstream-use documentation (intended purpose, permitted/restricted/prohibited uses, potential for deviation).
US
Failed
Developers of critical-impact AI systems who provide technology or services to deployers under contract or license must furnish deployers with (1) an overview of training data including size, sources, copyrighted data, and PII; (2) documentation of the baseline model's structure, input/output modality, size, and architecture; (3) known capabilities, limitations, and risks; and (4) downstream-use documentation including intended purpose, permitted/restricted/prohibited uses, and potential for deviation.
US
Failed
Covered entities must maintain and keep updated documentation of all data or input information used to develop, test, maintain, or update the automated decision system, including sourcing details, metadata, collection methodology, consumer consent status, data selection rationale, dataset representativeness, and data quality measures.