Document Provenance (Accessibility)
Field: Accessibility
In document accessibility, document provenance is the specific business process, template, and tools that produced a document, traced so that accessibility defects shared by many documents can be attributed to a common source and corrected there.
Scope
Document provenance concerns the workflow that produced a document’s defects, not the individual defects themselves. It is distinct from file-level remediation of a single document, such as adding tags, correcting reading order, or writing alternative text.
How provenance is traced
A file’s metadata often records the software that produced it. In PDF files, the Producer and Creator entries in the document information dictionary commonly name the application, library, and version, such as a scanner’s PDF library or an OCR tool. Grouping a collection of documents by these values shows which workflows are still producing documents with the same defects, including out-of-support tools and PDF versions.
Research on scholarly PDFs has found associations between the software used to create a document and whether it met accessibility criteria (Kumar and Wang, 2024).
Document provenance tracing is one analysis tool applied to a document corpus.
Distinctions
- Provenance in other fields. Provenance has established meanings in art history, archival science, and data lineage, where it refers to a record of origin and custody. The accessibility sense focuses on the production workflow as the cause of recurring defects.
History
The term appears in this sense in Document Provenance: Finding the Common Ancestor of Inaccessible Documents (Rietta, 2026).
Sources
- Uncovering the New Accessibility Crisis in Scholarly PDFs, Kumar and Wang, ACM SIGACCESS ASSETS '24. Retrieved September 23, 2026.
See also
Go deeper
Articles
- Document Provenance: Finding the Common Ancestor of Inaccessible Documents September 20, 2026