Document Provenance (Accessibility)

Field: Accessibility

In document accessibility, document provenance is the specific business process, template, and tools that produced a document, traced so that accessibility defects shared by many documents can be attributed to a common source and corrected there.

Scope

Document provenance concerns the workflow that produced a document’s defects, not the individual defects themselves. It is distinct from file-level remediation of a single document, such as adding tags, correcting reading order, or writing alternative text.

How provenance is traced

A file’s metadata often records the software that produced it. In PDF files, the Producer and Creator entries in the document information dictionary commonly name the application, library, and version, such as a scanner’s PDF library or an OCR tool. Grouping a collection of documents by these values shows which workflows are still producing documents with the same defects, including out-of-support tools and PDF versions.

Research on scholarly PDFs has found associations between the software used to create a document and whether it met accessibility criteria (Kumar and Wang, 2024).

Document provenance tracing is one analysis tool applied to a document corpus.

Distinctions

  • Provenance in other fields. Provenance has established meanings in art history, archival science, and data lineage, where it refers to a record of origin and custody. The accessibility sense focuses on the production workflow as the cause of recurring defects.

History

The term appears in this sense in Document Provenance: Finding the Common Ancestor of Inaccessible Documents (Rietta, 2026).

Sources

See also

Go deeper

Articles