Marton Ribary · Informatics 2020 · computational corpus linguistics study · n=?

A Corpus Approach to Roman Law Based on Justinian’s Digest

Cited 3 times in the scientific literature.

Level 4 - case-series / case-control

Level 4 by design analogy for computational corpus analysis of a single historical text collection (non-clinical scholarship).

OpenAlex W3092978792 · doi:10.3390/informatics7040044 · record verified 2026-08-31

What was done

Authors constructed a relational database of the Latin text of Justinian's *Digest* (533 CE) using Python. They applied computational clustering methods to analyze the text's linguistic profile and used computational distributional semantics to construct and compare Latin word embedding models for general and legal usage.

What was found

The abstract reports no numerical values, counts, or statistical metrics. Qualitatively, clustering identified an empirical structure inherent to the source texts rather than imposed abstract frameworks. Word embedding models detected a semantic split between general and legal senses of words, which the authors interpret as supporting a practical, practice-oriented focus in Roman legal education.

Why it matters

This work demonstrates how natural language processing and distributional semantics can be applied to ancient legal corpora to provide an empirical basis that complements traditional philological close reading.

Limits

The abstract provides no quantitative metrics, sample sizes (such as total word or section counts), or formal evaluation against expert philological standards. The study is restricted to a single compilation (*Justinian's Digest*), limiting generalizability across the broader corpus of ancient legal texts.

Cited by