Document Corpus (Accessibility)

Field: Accessibility ยท Also called: Document corpora, Corpus-level analysis

In document accessibility, a document corpus is the complete body of documents an organization publishes, treated as a unit of assessment in its own right. Some findings, such as the same non-descriptive title on every document or defects traced to a shared production workflow, exist only at the corpus level and cannot be seen by reviewing files one at a time.

Scope

The documents in a corpus are chiefly what the ADA Title II rule calls conventional electronic documents: “web content or content in mobile apps that is in the following electronic file formats: portable document formats (“PDF”), word processor file formats, presentation file formats, and spreadsheet file formats” (28 CFR ยง 35.104).

A document corpus is distinct from a website. Some web pages may belong to a corpus when they are published as documents in their own right, but the website as a whole, its navigation, templates, and applications, is not the corpus. Treating the two as the same would reduce the corpus to an inventory of URLs.

The test of when a web page might be considered a document is open for professional judgment based on a totality of the circumstances of how it functions. A major blog post or article might be a document whereas an index listing would not be. Web pages also have to meet ADA compliance requirements, though the remediation differs from most electronic documents.

Corpus-level analysis

Some findings are visible only when documents are examined together:

  • Repeated defects. A title stamped identically on every document by the same tool is present in each file, so a check of one file at a time can report a title as present while the corpus as a whole lacks descriptive titles.
  • Common sources. Grouping documents by the software recorded in their metadata, such as the PDF Producer and Creator entries, shows which workflows are producing the same defects. See document provenance.
  • Distributions. File format versions, out-of-support tools, and publication date ranges, which bear on eligibility for the rule’s exceptions for archived and preexisting documents.
  • Completeness. Title II exceptions are determined document by document, so showing that every document was accounted for requires an inventory of the whole corpus.

Distinctions

  • Corpus linguistics. Corpus analysis is an established method in linguistics, the study of language through large collections of text. The accessibility sense concerns the documents’ conformance and production, not their language.
  • Information retrieval and machine learning. A corpus is a collection of documents indexed for search or used as a dataset. Kumar and Wang (2024) analyzed a corpus of scholarly PDFs in this sense.
  • E-discovery. In litigation, the documents collected for review are sometimes called the corpus or the review set.

History

The term appears in this sense in Threat Modeling for ADA/WCAG Compliance (Rietta, July 2026).

Sources

See also

Go deeper

Articles