Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States | Summary
17 Jun 2026 | Paper Review Legal NLP Datasets LLM-as-a-JudgeContents
- Summary
- 1 What it means to “free the law”
- 2 Related Work
- 3 Properties of LOCUS
- 4 Constructing LOCUS
- 5 A Dimensional Analysis of Local Laws
- 6 Discussion, Limitations, and Future Work
- Appendix
- Brief Thoughts
This article explains the key points of Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States.
- 2026-06-17 (arXiv)
- Peskoff, Denis, Barrow, Joe, Vu, Christopher, Davenport, Diag.
- UC Berkeley, School of Information, Independent
- Paper
Summary
- Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States introduces a corpus of codes from 9,239 cities and counties, addressing the fragmentation of local ordinances across platforms built for browsing rather than bulk research access. LOCUS-v1 releases a county-harmonized layer with 2,211,516 text chunks covering 2,309 of 3,144 U.S. counties; those counties represent 94% of the population, but the selected codes do not necessarily apply to every resident within them.
- The pipeline converts roughly 7M PDF pages to Markdown, removes formatting artifacts, segments individual laws, and applies ModernBERT classifiers trained on LLM-generated annotations. Separate regressors learn enforcement discretion, opacity, paternalism, and problem salience from pairwise judgments aggregated with TrueSkill.
- On 1,000 held-out ordinances per dimension, the regressors achieve Pearson correlations of 0.822–0.936 with normalized TrueSkill targets. These results measure agreement with LLM-derived scores, not independent legal validity; the release leaves issue-specific jurisdictional hierarchy and controlling-authority reasoning unresolved. The dataset and derivative models are available at https://huggingface.co/datasets/LocalLaws/LOCUS-v1.
1 What it means to “free the law”
Local ordinances govern everyday activities such as zoning, housing, licensing, noise, and animal control, yet existing legal AI resources lack systematic access to this layer of law. LOCUS provides a national corpus while recognizing that retrieval alone cannot resolve interactions among state statutes, county ordinances, municipal codes, charters, home-rule provisions, and preemption doctrines.
- Vendor-specific navigation, dynamically generated PDFs, inconsistent jurisdiction indexes, and the absence of a central hosting registry make national collection a discovery and extraction problem rather than a single scraping task.
- The county-harmonized layer selects a representative code by document length, providing a common geography for search and links to Census and policy data without deciding which authority controls a legal question. Retrieval, structured regulatory extraction, and jurisdiction-sensitive reasoning benchmarks are proposed applications, not evaluated systems.
- Figure 1 distinguishes county-code selections, city-code selections, and uncovered counties. Figure 2 illustrates ordinance-level metadata: the festival rule for minors receives high predicted paternalism (+5.48) and low opacity (−2.54), whereas the false-emergency-alarm prohibition receives higher problem salience (+2.62).
The map distinguishes the source selected for each covered county: a county code or a city code, with a separate category for no coverage. It visualizes source selection and geographic gaps, not the territorial applicability of every selected ordinance. This distinction is necessary when interpreting the reported 94% population statistic.
The false-emergency-alarm prohibition from Paris Municipal Code and the festival-attendance rule from Mecklenburg County Code both receive Rules and Nuisance labels, but their dimensional profiles differ. The festival rule has predicted paternalism +5.48 and opacity −2.54, while the false-alarm rule has problem salience +2.62. Identical categorical labels can therefore coexist with substantially different continuous scores.
2 Related Work
The paper positions LOCUS alongside legal NLP corpora covering case law, opinions, contracts, and statutes, including ECHR and Pile of Law. It notes that none of LegalBench’s 162 tasks involve local ordinances, distinguishing corpus construction from the future development of local-law reasoning benchmarks.
- Georgia v. Public.Resource.Org, Inc., No. 18-1150, decided April 27, 2020, supplies the paper’s legal-access context. Public-domain status does not itself eliminate the technical barriers to collecting and standardizing local codes.
3 Properties of LOCUS
LOCUS separates the broad raw collection from a smaller county-harmonized release intended for retrieval and comparative analysis. Structured annotations describe the function and topic of text chunks, while additional documents are intended for controlled researcher access.
3.1 A County-Harmonized Access Layer
The public access layer contains 2,211,516 chunks, with Rules and Enforcement designated as substantive and Context and Process retained as non-substantive text. Structural artifacts are removed; substantive chunks receive one of five merged topic labels: Buildings, Business, Zoning, Nuisance, or Other.
- Figure 3 shows Rules as the largest function category. Its rounded counts should not be treated as exact totals for the release.
- Roughly a third of substantive laws fall under Other, which the authors report identifying with near 90% precision. Inspection of headers associates this category with government, employment matters, and animal regulation; the latter contributes to Alaska’s comparatively large share of Other chunks.
- The county representation selects the longer available code from the county and a municipality, ideally its largest city. This is a reproducible access convention, not a legal determination that the selected text governs the entire county.
The function chart shows Rules dominating the retained corpus, with rounded labels of 1.4M Rules, 300k Enforcement, 300k Context, and 200k Process. The topic chart covers only Rules and Enforcement and shows Other as its largest category at 600k. These rounded labels summarize the distribution rather than supplying an exact accounting of all 2,211,516 chunks.
3.2 Additional Data for Researchers
Beyond the public release, the authors collected an additional 7,000 documents from other cities and counties. They intend to provide researcher access through signed releases, drawing an analogy to MIMIC, to help preserve future evaluations of foundational models’ local-law coverage against ingestion contamination.
4 Constructing LOCUS
Figure 4 summarizes the construction workflow: collect PDFs, perform OCR to Markdown, join and clean page-level output, segment laws, and classify and score each segment. The two final branches provide categorical organization of legal text and continuous measurement of four normative dimensions.
The workflow standardizes more than 9,000 PDFs and 7M pages through OCR, cleanup, and segmentation. It then branches into categorical classification and four-dimensional scoring, showing that the same extracted law supports both organizational metadata and comparative measurement.
4.1 Collecting the data
The raw collection contains 9,239 valid PDFs totaling approximately 80 GB. Collection combines browser automation and vendor-specific download logic with manual recovery of self-hosted or PDF-restricted codes that automated workflows do not cover.
- Observed failures include server-side PDF assembly limits, filename collisions between municipalities with identical names, hidden interface thresholds, 15 second crawl delays, anti-bot measures, and consolidated cities spanning multiple counties.
- These failures required targeted recovery rather than a generic scraper, making collection engineering a substantive part of the contribution.
4.2 Identifying salient laws
The initial labeling strategy uses zero-shot LLM classification to distinguish substantive laws from structural and procedural content. After comparing GPT-5.4 mini and nano on a 500-sample evaluation, the authors select nano as the lower-cost annotator and review the 5.5% of annotations deemed most challenging with GPT-5.4.
- The stronger model agrees with 64,977 of 108,889 reviewed predictions. Disagreements often shift Rules labels toward Process or Enforcement, indicating uncertainty in the functional boundaries of the annotation scheme.
- The authors report that structural content was consistently recognized and removed. They identify evaluation by lawyers and judges as a way to move beyond the limitations of LLM-as-a-Judge.
4.3 OCR and Processing
LightOnOCR-2-1B converts every page image into Markdown to accommodate scanned, born-digital, single-column, and double-column documents. Post-processing removes repeated headers, footers, and page numbers, merges paragraphs and tables across pages, and identifies section and subsection headers for segmentation.
- The authors describe LightOnOCR-2-1B as an open 1B parameter vision-language model fine-tuned on 16MM PDF pages. They report robust reading order across ordinance formats, but provide no corpus-specific quantitative OCR accuracy evaluation.
- The roughly 7M-page workload runs on Modal, https://modal.com, at approximately $0.30 per 1,000 pages. This is the reported processing rate, not a complete budget for collection, annotation, and model training.
- ModernBERT-base classifiers then infer substantivity, function, and topic for extracted segments; segments classified as purely Structural are omitted.
4.4 Annotating the Law
Three classifiers organize ordinances by substantivity, function, and topic. GPT-5.4-nano annotates a sample of 100,000 laws, split into 80,000 training instances, 10,000 instances for parameter sweeps, and a final 10,000-instance evaluation subset.
- Table 1 distinguishes operative Rules, Enforcement provisions, explanatory Context, administrative Process, and non-operative Structural artifacts. Topic examples are restricted to texts labeled Rules or Enforcement.
- The five merged topic labels include heterogeneous residual content under Other. The appendix’s initial annotation prompt uses a finer seven-category topic vocabulary, which should not be conflated with the five merged release labels.
- The paper specifies the evaluation split but does not report a complete set of held-out performance metrics for the three classifiers.
The examples clarify the label boundaries: a direct-sales permit requirement is Rules, officer powers are Enforcement, a section listing is Structural, stated procedural purpose is Context, and license-application steps are Process. Topic examples demonstrate the five merged categories, including budget-allocation restrictions under Other. This residual category contains operative legal material rather than only unclassified noise.
4.5 Creating a Harmonized Access Layer
For each county, the harmonization algorithm checks for a county code and an available city code, ideally from the largest city, and selects the longer document by page count when both exist. The covered counties represent 94% of the U.S. population, but the authors distinguish this geographic statistic from the smaller population literally governed by the selected codes.
- Page length is chosen for interpretability and reproducibility, supported by an observed association between code length and jurisdiction population.
- County codes are slightly shorter on average than city codes, but the algorithm introduces no jurisdiction-specific weights. Other municipalities and residents outside a selected city remain incompletely represented by the access layer.
5 A Dimensional Analysis of Local Laws
LOCUS assigns continuous scores for Enforcement Discretion, Opacity, Paternalism, and Problem Salience, enabling comparison within a code and across jurisdictions. The dimensions concern official latitude in enforcement, difficulty understanding obligations, protection of the actor versus protection of others, and textual emphasis on an issue’s importance or threat.
- Continuous scoring is intended to order laws rather than assign them to discrete normative classes. Figure 2 combines these scores with function and topic labels at the ordinance level.
- Figure 5 maps length-residualized opacity and paternalism, separating geographic patterns on the two axes rather than treating them as a single measure of regulation.
The maps display length-residualized opacity and paternalism on separate scales, with endpoints −0.36 and +0.36 for opacity and −0.29 and +0.29 for paternalism. Florida appears relatively opaque but comparatively non-paternalistic, illustrating why the two axes should not be collapsed into a single regulatory score. The maps describe model-derived textual variation rather than legal authority or enforcement outcomes.
5.1 Building LOCUS Scorers
For each dimension, GPT-5.4-nano produces 200,000 pairwise comparisons over a fixed sample of 10,000 ordinances, choosing A, B, or Tie. Each comparison pair is also judged in reverse order to address position bias, and TrueSkill aggregates match histories into latent scores.
- The TrueSkill scores are normalized by subtracting the dimension’s mean and dividing by its standard deviation.
- Each dimension uses 8,000 training, 1,000 validation, and 1,000 test ordinances. A separate ModernBERT-base model with a linear regression head learns the normalized scores using mean-squared error.
Figure 6 reports test-set Pearson correlations of 0.822 for Paternalism, 0.909 for Opacity, 0.872 for Enforcement Discretion, and 0.936 for Problem Salience. The regressors reproduce much of the variation in LLM-derived TrueSkill targets, with paternalism the least closely reproduced dimension.
- Each correlation is measured on 1,000 held-out ordinances for its dimension. The targets remain products of pairwise LLM judgments rather than independently established human scores.
- The paper’s score-viewing website is https://locallaws–locus-leaderboards-web.modal.run.
The scatterplots compare ModernBERT predictions with normalized TrueSkill targets on 1,000 test ordinances per dimension. Correlations are 0.822 for paternalism, 0.909 for opacity, 0.872 for enforcement discretion, and 0.936 for problem salience. Their alignment supports the regressors as approximations to the pairwise-judgment pipeline, but does not independently validate the underlying normative judgments.
5.2 Analysis
The authors report that county codes are more opaque than city codes on average and that Florida’s opacity is more than twice that of any other state. Across 2,211,516 sections, opacity and paternalism are weakly correlated, with Pearson r=0.11.
- Figure 5 shows Florida as relatively opaque but not paternalistic in the length-residualized maps. These are model-based descriptive comparisons, not evidence about causal effects or actual enforcement behavior; standardized opacity scores also do not establish a ratio-scale measure of legal difficulty.
- High paternalism scores help identify curfews and sections whose headers contain ‘possession’ or ‘alcoholic’; ‘definitions’ and ‘variances’ are associated with opacity.
6 Discussion, Limitations, and Future Work
LOCUS-v1 is an access layer rather than a complete representation of legal authority: a representative code cannot establish the controlling rule for a person, parcel, business, or issue. The discussion emphasizes preserving jurisdiction type and document structure because city and county codes differ substantively, and those differences vary by region.
- The authors report more zoning material in county codes and more nuisance and public-order regulation in city codes. In the Northeast, counties appear less zoning-heavy and more enforcement-oriented, limiting uniform interpretations of county-level harmonization.
- Codes exhibit a recurring topic sequence: general provisions and governmental structure, business regulation, nuisance and public order, zoning, and building regulation. The discussion presents this as a reason to retain code position in retrieval and benchmark design, though it supplies no detailed quantitative evaluation of the sequence.
- Future benchmarks would need to identify relevant government layers, incorporate state-law context, handle overlapping sources, and determine controlling authority. These capabilities are proposed uses of LOCUS, not demonstrated system results.
Appendix
- A Scoring Prompts provides the Pairwise Comparison System Prompt and four axis rubrics. Axis Rubric: Problem Salience concerns urgency, severity, and heightened penalties signaling gravity; Axis Rubric: Paternalism vs. Externality Orientation distinguishes self-regarding harms from harms to others; Axis Rubric: Opacity / Intelligibility focuses on jargon, cross-references, undefined terms, and convoluted structure; Axis Rubric: Enforcement Discretion combines breadth of citizen exposure with officials’ textual latitude, reserving the floor for provisions that do not act on private parties. The comparison template requests a winner of A, B, or Tie and a one-sentence explanation.
- B Annotation Prompt includes the Annotation Prompt, review instructions, and CLASSIFICATION_SCHEMA for initial GPT-5.4-nano classification and GPT-5.4 review. The schema requires is_substantive, primary_function, sub_category, and logic; is_substantive is 1 only for Rules or Enforcement, with sub_category otherwise null. The initial topic vocabulary is Land use, Noise/Nuisance, Housing, Business licensing, Public space, Building/Safety, and Other. Review guidance distinguishes operative internal governance procedures labeled Process from non-operative artifacts labeled Structural and requests confirm, override, or fresh outcomes with review_logic, although these review fields are absent from the printed schema.
Brief Thoughts
The main contribution is the conversion of fragmented local codes into searchable, geographically organized infrastructure at a documented scale of 9,239 PDFs and 2,211,516 released chunks. Its value depends on retaining the distinction between geographic coverage and governing authority. The scorer correlations support efficient reproduction of LLM-derived rankings, but do not establish agreement with expert legal judgments. Substantial disagreement on difficult function labels reinforces the need for human review. Comparative analysis should retain municipal versus county provenance and account for the longest-code selection rule. Corpus-specific OCR validation, fuller classifier metrics, and expert assessment would strengthen confidence in fine-grained legal conclusions.