AI hiring compliance across six regimes
Clarion is a compliance overlay that sits on top of the AI hiring tools an employer already runs. It reads one vendor scoring export, computes every adverse-impact statistic in deterministic code, and fans that single audit into six deliverables shaped to NYC, Colorado, Illinois, Texas, California and the EU. Where two of those regimes demand opposite things from the same feature, it says so and quantifies both exposures rather than printing "fully compliant." What follows is a demo over a seeded synthetic export, not a deployed pipeline.
6
Jurisdiction deliverables from one audit run
NYC, Colorado, Illinois, Texas, California, EU
9 of 13
Obligations routed to a named human, not auto-cleared
7 NEEDS PROOF plus 2 CONFLICT, on the demo's seeded export
1
Conflict the gate refuses to bluff past
Illinois HB 3773 against EU AI Act Article 10(3)
The dataset is synthetic: a seeded export from "Acme Logistics, Inc." of 1,040 candidates on requisition REQ-2026-0412, scored by three simulated tools. The Workday-Spotlight-style, HireVue-style and Eightfold-style scorers are archetypes served by stubbed fixture adapters, not integrations, and no real employer or candidate appears anywhere in it.
A CHRO, General Counsel or Chief Risk Officer running automated hiring tools across two or more of NYC, Colorado, Illinois, Texas, California and the EU answers to six separate regimes with one stack. The vendors serving that stack hand over a single "we passed the four-fifths rule" audit as if it settles all six. It does not. NYC Local Law 144 asks for intersectional impact ratios. Colorado SB 24-205 asks for a documented reasonable-care program and specifies no methodology at all. Texas TRAIGA rejects disparate impact as a standalone basis and asks about intent, which makes the exact statistics driving the NYC deliverable evidentiarily irrelevant to the Texas one. The EU AI Act asks whether your training data is representative, and leans on the very geographic feature Illinois prohibits as a proxy.
Researchers who went looking for published Local Law 144 bias audits found them for 4.6% of 391 NYC employers, the finding they named Null Compliance (Cornell / Data & Society / Consumer Reports, FAccT 2024). Publishing the audit is the visible half of the obligation, and most of that sample had not reached it.
The NY State Comptroller found 17 potential Local Law 144 violations in the same 32-company sample where DCWP had found one, and DCWP agreed to shift to proactive enforcement (NY State Comptroller, December 2, 2025). The gap between a passing vendor audit and a defensible record is where that enforcement lands.
Scope self-classification, ADA accessibility on video interviews and FCRA adverse-action duties are separate legal theories a passed bias audit never tested. Mobley v. Workday, D.K. v. Intuit/HireVue and Kistler v. Eightfold each raise one of them, and none is decided.
The pipeline is short and the split of work inside it is strict. A vendor AEDT export is normalized into one canonical record schema, a deterministic engine computes every statistic, six regime rule packs and a policy gate decide every obligation, an agent crew writes the prose, and the run seals a hash-chained pre-audit package. Agents advise, code decides, which is what keeps the output filable regardless of how good or bad the model underneath is.
01 / THE DETERMINISTIC TRUST CORE
A numpy engine computes marginal four-fifths impact ratios, intersectional race by sex ratios with Benjamini-Hochberg FDR control, protected-class proxy detection by Cramér's V, a counterfactual protected-attribute flip through a transparent logistic-regression surrogate, an ASR word-error-rate disparity on the video subset, and an FCRA trigger predicate. Same input, same output, every run.
02 / SIX RULE PACKS AND THE POLICY GATE
Each regime carries its own citation, effective date and required deliverable shape, and each obligation resolves to PASS, FAIL, NEEDS PROOF or CONFLICT. The gate is the one thing the language model cannot reach, so a well-argued narration can never move an obligation from NEEDS PROOF to green.
03 / THE AGENT CREW
Six jurisdiction agents narrate each regime's verdict, a conflict reconciler turns a hard conflict into a legal-strategy memo with the exposure of each branch, and an adversarial skeptic tries to refute every claimed pass and routes what it cannot clear to human proof. The crew is provider-swappable; with no provider configured it falls back to deterministic templates and the app runs identically.
04 / THE PRE-AUDIT PACKAGE
The run seals a SHA-256 hash-chained bundle of 17 nodes, each carrying its inputs, its computation and its rule citation plus a link to the hash of the node before it, so any edit breaks the chain. It exports as JSON and as a printable HTML auditor packet, and the chain is verified on every run and unit-tested for tamper-evidence.
The console is two views plus click-through dialogs rather than a single dashboard. Manual Testing holds the stage bar, the live Audit Execution trace and the candidate board. Run Benchmark, which stays disabled until Run Audit finishes, switches to Benchmark Results: three tiles reading Jurisdiction Deliverables 6, Lowest Impact Ratio 0.65 and Items Requiring Human Action 9, the six Jurisdiction Deliverables rows, a Supporting Evidence button row and Export Pre-Audit Package in the header.
Everything underneath opens as a dialog. Selection-Rate Analysis, the Conflict Register and the Human Proof Queue are the three Supporting Evidence buttons, and each regime's obligations and audit narrative open from its own deliverable row. Nothing is summarized into a single score, because a single score cannot be true across six differently shaped questions.
The demo audits a seeded synthetic export of 1,040 candidates on requisition REQ-2026-0412, scored by three simulated vendor tools. The violation in it is planted and reproducible, built so the engine has something real to catch. Here is what the run puts on screen, in the order it puts it there.







Every figure on this page describes one seeded synthetic export of 1,040 candidates with a planted, reproducible violation. They demonstrate that the engine catches what a marginal-only audit misses. They are not an accuracy rate, not a benchmark against other products, and not a claim about any real employer's hiring data.
| Question | What Clarion does in this demo | What remains outside the demo |
|---|---|---|
| Coverage | One audit run emits six jurisdiction-shaped deliverables across 13 obligations, with 2 PASS, 2 FAIL, 7 NEEDS PROOF and 2 CONFLICT. | A compliance certificate or a fairness score. The gate is built to refuse one. |
| Adverse impact | Marginal and intersectional four-fifths ratios with Benjamini-Hochberg FDR control, proxy detection by Cramér's V, and a counterfactual flip moving advance probability by 0.0574 on average and 0.1228 at maximum. | Any statement about real hiring outcomes. The 1,040 candidates, the requisition and the disparity are synthetic and seeded. |
| Vendor data | Reads one normalized AEDT export produced by Workday-Spotlight-style, HireVue-style and Eightfold-style scorers, served by stubbed fixture adapters. | Live connectors into any ATS, assessment platform or match engine. No named vendor is a customer, partner or endorser. |
| Downstream action | Detects and routes the ADA and ASR accessibility finding and the FCRA adverse-action trigger to a named human, with the evidence attached. | The CART accommodation workflow and the candidate-facing dispute portal. Both are simulated in the demo and are not something we build here. |
| Sign-off | Seals a 17-node hash-chained pre-audit package, verified on the run and unit-tested for tamper-evidence, in JSON and printable HTML. | The independent bias audit itself. Firms like DCI Consulting, ORCAA and Secretariat hold that role, and we do not. |
Clarion does not certify compliance, and refusing to is the design. It is not a hiring model and never scores or ranks a candidate. The dataset is synthetic and seeded, the Black / Female disparity in it is planted so the engine has something real to catch rather than a finding about any actual employer, and the vendor connectors behind it are stubbed fixture adapters rather than live integrations. The tool produces evidence for counsel: it does not give legal advice or opine on liability, and Mobley v. Workday, D.K. v. Intuit/HireVue and Kistler v. Eightfold are pending theories rather than decided outcomes. Statutory maxima are quoted only as maxima, and no modeled exposure figure appears here. This page is an explainer with the walkthrough video, real screenshots, the mechanism and answers, not an app you operate from here.
Because the test that passes and the test the law asks for are not the same test. On the demo's seeded 1,040-candidate synthetic export the marginal four-fifths check passes, with a minimum race impact ratio of 0.8196 and a minimum sex ratio of 0.8744 against the 0.80 line, and that is the audit a vendor self-report ships. NYC Local Law 144 asks for the intersectional race by sex ratios, and there the Black / Female cell advances 44 of 130 at an impact ratio of 0.6471. A marginal-only audit is not wrong about that cell, it is structurally blind to it.
No. Veriprajna is not the independent auditor, and that role belongs to firms like DCI Consulting, ORCAA or Secretariat. Clarion produces the pre-audit package, a hash-chained bundle in which every number carries its computation and every verdict carries its citation, built so an independent auditor can sign it without rewriting it. The goal is to get the system into the state where the audit finds nothing worth writing up.
It connects to neither, and Clarion never scores or ranks a candidate itself. Nothing in your stack is replaced. It reads a vendor AEDT export and normalizes it into one canonical record schema. In this demo the vendor connectors are stubbed fixture adapters over synthetic data, and the Workday-Spotlight-style scorer, the HireVue-style video round and the Eightfold-style match engine are archetypes rather than integrations.
Every statistic, every threshold comparison, every conflict edge and the bundle hashing sit in deterministic code outside the agent framework. The six jurisdiction agents, the conflict reconciler and the adversarial skeptic write prose; they never compute a number and they cannot override the policy gate. With no model provider configured the crew falls back to deterministic templates and the app produces the same verdicts, because the verdicts were never the model's to give.
You get the trade-off quantified and written down. In the demo the zip_region feature correlates with race at a Cramér's V of 0.3321, over the 0.2 threshold, so Illinois HB 3773 treats it as a banned proxy while EU AI Act Article 10(3) representativeness leans on the same geographic coverage. The gate emits CONFLICT instead of a pass, and the reconciler writes a strategy memo: run two deployment configurations, or accept one exposure knowingly and document it, noting that EU high-risk penalties reach a statutory maximum of the greater of 15 million euros or 3% of global annual turnover.
It automates the part that is honest to automate. On the demo's synthetic export 2 of 13 obligations are auto-satisfied, 2 fail outright, and 9 are routed as NEEDS PROOF or CONFLICT to a named human with the evidence already assembled and the citation attached. No honest system clears all thirteen on this data. A dashboard that showed thirteen greens would be clearing them without proof, and that dashboard is the exhibit a plaintiff reads back to you. The engine also clears school_tier as not a proxy at a Cramér's V of 0.0711, so it is not flagging everything it sees.
One export produces a pre-audit package as JSON and as a printable HTML auditor packet, version 1 of the veriprajna-pre-audit-package format, carrying 17 nodes with the chain verified on the run. Each node holds its inputs, its computation and its rule citation plus a SHA-256 link to the node before it, so any edit to any node breaks the chain. Alongside it sit the six jurisdiction deliverables, each shaped to its own regime's citation, effective date and required format.
The research behind this demo — the architecture, the verification design, and the enterprise blueprint.
We are an AI engineering team, not a compliance certifier. We build the deterministic layer that computes the statistics, shapes the deliverable each regulator asked for, and names the contradictions out loud so counsel can decide with the exposure in front of them.
A useful first conversation is concrete: which AEDT vendors sit in your funnel, which of the six regimes you are actually exposed to, and which features in those models would look like protected-class proxies under one of them. We can work through the engine, the rule packs and the pre-audit package format alongside your HR, legal and data teams.