CA
Enacted eff 2026-01-01
Developers must post on their website, by January 1, 2026, and before each subsequent public release or substantial modification, documentation of the training data used for any generative AI system or service released on or after January 1, 2022, including dataset sources, data types, volume, IP status, personal information presence, processing methods, collection timeframes, and use of synthetic data. Exemptions apply for systems whose sole purpose is security and integrity, national airspace operations, or national security/military/defense systems available only to federal entities.
NY
Engrossed
Developers must post on their website, on or before January 1, 2027, and before each subsequent public release or substantial modification of a generative AI model or service released on or after January 1, 2022, documentation regarding the data used to train the model or service, including a high-level summary of the datasets covering: (1) the sources or owners of the datasets; (2) a description of how the datasets further the intended purpose; (3) the number of data points, which may be in general ranges with estimates for dynamic datasets; (4) a description of the types of data points (for labeled datasets, the types of labels used; for unlabeled datasets, general characteristics); (5) whether the datasets include data protected by copyright, trademark, or patent or are entirely in the public domain; (6) whether the datasets were purchased or licensed; (7) whether the datasets include personal information or personal identifying information; (8) whether the datasets include aggregate consumer information; (9) whether there was any cleaning, processing, or other modification, including its intended purpose; (10) the time period during which data was collected, including notice if collection is ongoing; (11) the dates datasets were first used during development; and (12) whether the model uses or continuously uses synthetic data generation. This obligation does not apply to models solely for national airspace aircraft operations or models developed for national security, military, or defense purposes available only to a federal entity.
GA
Introduced eff 2027-01-01
Production companies deploying AI systems for use in production in Georgia must issue a disclaimer describing how AI was adopted and deployed and identifying any data, sources, or metrics that were used.
NY
Introduced
Developers must post on their website, before January 1, 2027 and before each subsequent public release or substantial modification of a generative AI system released on or after January 1, 2022, detailed information about journalism content used in training: (1) URLs accessed by crawlers, (2) a description of the content used sufficient to identify individual works, including type, provenance, and means of acquisition, (3) whether source identifiers, terms, or copyright notices were removed, and (4) the timeframe of data collection. This obligation is excused where an express written agreement authorizes the developer's access and both parties agree not to post the information.
NY
Introduced
Developers who deploy crawlers must publicly disclose, in a manner clearly accessible to website operators and updated in real time with any changes, the identity and technical details of each crawler: (1) crawler name, IP address, and user-agent identifier, (2) the legal entity responsible, (3) the specific purposes for which each crawler is used, (4) legal entities receiving scraped data, and (5) a single point of contact for third-party communications and complaints.
NY
Introduced
Developers must post on their website, on or before January 1, 2027, and before each subsequent public release or substantial modification of a generative AI model or service released on or after January 1, 2022 and made available to New Yorkers, documentation regarding the data used to train the model or service. The documentation must include a high-level summary of the training datasets covering at least twelve categories: (1) sources or owners of the datasets; (2) how the datasets further the model's intended purpose; (3) the number of data points (in general ranges, with estimates for dynamic datasets); (4) the types of data points (label types for labeled datasets, general characteristics for unlabeled datasets); (5) whether the datasets include data protected by copyright, trademark, or patent, or are entirely in the public domain; (6) whether the datasets were purchased or licensed; (7) whether the datasets include personal information or personal identifying information; (8) whether the datasets include aggregate consumer information; (9) whether there was any cleaning, processing, or other modification, and the purpose of those efforts; (10) the time period during which data was collected, including notice if collection is ongoing; (11) the dates datasets were first used in development; and (12) whether the model uses or continuously uses synthetic data generation. Exemptions apply for generative AI models whose sole purpose is aircraft operation in the national airspace, and for models developed for national security, military, or defense purposes available only to a federal entity.
NY
Introduced
Developers must post on their website, by January 1, 2027 and before each subsequent public release or substantial modification, detailed information about journalism content used for AI utilization — including URLs accessed, content descriptions sufficient to identify individual works, whether source identifiers or copyright notices were removed, and data collection timeframes. This obligation does not apply where the developer has an express written agreement with the journalism provider authorizing content access and both parties agree not to post the information.
NY
Introduced
Developers deploying crawlers must, by January 1, 2027, publicly disclose on an easily accessible platform the identity of each crawler (name, IP address, user-agent identifiers), the responsible legal entity, specific purposes, downstream data recipients, and a single point of contact for complaints. This information must be kept current.
US
Introduced
Covered entities must disclose training data sources, documentation, testing methodology, data collection during inference, and operational information for each foundation model, both before commercial deployment and throughout the system lifecycle.
US
Introduced
Covered entities must disclose a sufficiently detailed summary of training data sources, collection methods, inference-time data collection and retention practices, and the size and composition of training data including demographic, language, and attribute information, while accounting for privacy.
US
Introduced
Covered entities whose foundation model is derived from or built upon another covered entity's foundation model must independently comply with all transparency regulations for any significant change, retraining, or adaptation they apply to the base model.
VA
Introduced
Developers must post on their website, within 72 hours of making a generative AI system available in Virginia and within 72 hours after each significant update, detailed documentation for each training dataset covering: dataset name, source, volume, IP status, Do Not Train data handling, personal data management, illegal material screening, collection timeframe, and synthetic data generation use. Systems whose sole purpose is security or integrity are exempt.
WA
Introduced
Developers must post on their website, by January 1, 2026 and before each subsequent public release or substantial modification, documentation regarding the data used to train any generative AI system or service released on or after January 1, 2022 and made publicly available to Washington residents. The documentation must include a high-level summary of training datasets covering: (1) the sources or owners of the datasets; (2) how the datasets further the system's intended purpose; (3) the number of data points (general ranges and estimates permitted for dynamic datasets); (4) the types of data points within the datasets; (5) whether the datasets were purchased, licensed, or publicly available; (6) whether the datasets include personal information as defined in RCW 19.373.010; (7) whether the datasets include aggregate consumer information; (8) whether any cleaning, processing, or modification was performed and its intended purpose; (9) the dates the datasets were first trained or last significantly updated; and (10) whether the system used or continuously uses synthetic data generation. This obligation does not apply to systems whose sole purpose is security and integrity, aircraft operation in the national airspace, or national security/military/defense purposes available only to a federal entity.
WA
Introduced eff 2027-01-01
Developers must post on their website, by January 1, 2027 and before each subsequent public release or substantial modification, documentation of the datasets used to train each generative AI system released on or after January 1, 2022 — including dataset sources, data types, volume, IP status, personal-information presence, acquisition method, processing steps, training dates, synthetic-data use, and CSAM removal steps. Exempt: security-only systems, aviation systems, national-security/defense systems available only to federal entities, and FDA-regulated systems.
MD
Failed
Developers must publish on their website, by January 1, 2026, and before each subsequent release or substantial modification, documentation detailing the data and datasets used to train the generative AI system — including data sources, purpose alignment, dataset size, label types, IP status, licensing status, presence of personal information and aggregate consumer information, processing methods, collection timeframes, first use during development, and synthetic data generation.
US
Failed
Covered entities must disclose training data sources (including personal data collection and information to assist copyright owners), data size and composition (including demographics and language information), data governance procedures (including editing and filtering), and data labeling methods and validation, as specified in FTC regulations.