Paper Type
Complete
Abstract
Open dataset repositories function as critical infrastructure in the artificial intelligence data supply chain. Emerging governance expectations, including the EU AI Act, presuppose provenance documentation that is reviewable at scale. We translate these expectations into auditable indicators capturing accountable stewardship, origin statements, and traceability anchors, and audit 29,000 dataset descriptions from Hugging Face and Zenodo. Using a hybrid pipeline (manual annotation, embeddings, supervised classification, and targeted human review), we quantify component prevalence and end-to-end completeness via a strict complete provenance chain benchmark. Across platforms, individual signals appear intermittently, while complete provenance chains remain exceptionally rare. Within Hugging Face, disclosure is higher for datasets linked to downstream model training, yet complete chains remain uncommon even in this exposed subset. Regression analyses show provenance completeness aligns more with platform attention than with downstream exposure. Overall, voluntary documentation does not reliably yield the compliance-oriented evidence needed for accountable data governance.
Paper Number
1919
Recommended Citation
Sivizaca Conde, Daniel Juan; Rößler-von Saß, David; Perez, Sergio; and Kliewer, Natalia, "The Operationalization Gap: Auditing Provenance Transparency under the AI Act" (2026). AMCIS 2026 Proceedings. 44.
https://aisel.aisnet.org/amcis2026/sig_sec/sig_sec/44
The Operationalization Gap: Auditing Provenance Transparency under the AI Act
Open dataset repositories function as critical infrastructure in the artificial intelligence data supply chain. Emerging governance expectations, including the EU AI Act, presuppose provenance documentation that is reviewable at scale. We translate these expectations into auditable indicators capturing accountable stewardship, origin statements, and traceability anchors, and audit 29,000 dataset descriptions from Hugging Face and Zenodo. Using a hybrid pipeline (manual annotation, embeddings, supervised classification, and targeted human review), we quantify component prevalence and end-to-end completeness via a strict complete provenance chain benchmark. Across platforms, individual signals appear intermittently, while complete provenance chains remain exceptionally rare. Within Hugging Face, disclosure is higher for datasets linked to downstream model training, yet complete chains remain uncommon even in this exposed subset. Regression analyses show provenance completeness aligns more with platform attention than with downstream exposure. Overall, voluntary documentation does not reliably yield the compliance-oriented evidence needed for accountable data governance.
When commenting on articles, please be friendly, welcoming, respectful and abide by the AIS eLibrary Discussion Thread Code of Conduct posted here.
Comments
SIG SEC