AI hiring compliance across six regimes

One audit run, six jurisdiction-shaped deliverables, and a place where two of those regimes cannot both be satisfied.

Clarion is a compliance overlay that sits on top of the AI hiring tools an employer already runs. It reads one vendor scoring export, computes every adverse-impact statistic in deterministic code, and fans that single audit into six deliverables shaped to NYC, Colorado, Illinois, Texas, California and the EU. Where two of those regimes demand opposite things from the same feature, it says so and quantifies both exposures rather than printing "fully compliant." What follows is a demo over a seeded synthetic export, not a deployed pipeline.

6

Jurisdiction deliverables from one audit run

NYC, Colorado, Illinois, Texas, California, EU

9 of 13

Obligations routed to a named human, not auto-cleared

7 NEEDS PROOF plus 2 CONFLICT, on the demo's seeded export

1

Conflict the gate refuses to bluff past

Illinois HB 3773 against EU AI Act Article 10(3)

The dataset is synthetic: a seeded export from "Acme Logistics, Inc." of 1,040 candidates on requisition REQ-2026-0412, scored by three simulated tools. The Workday-Spotlight-style, HireVue-style and Eightfold-style scorers are archetypes served by stubbed fixture adapters, not integrations, and no real employer or candidate appears anywhere in it.

Six regulators looked at the same hiring stack and asked six structurally different questions.

A CHRO, General Counsel or Chief Risk Officer running automated hiring tools across two or more of NYC, Colorado, Illinois, Texas, California and the EU answers to six separate regimes with one stack. The vendors serving that stack hand over a single "we passed the four-fifths rule" audit as if it settles all six. It does not. NYC Local Law 144 asks for intersectional impact ratios. Colorado SB 24-205 asks for a documented reasonable-care program and specifies no methodology at all. Texas TRAIGA rejects disparate impact as a standalone basis and asks about intent, which makes the exact statistics driving the NYC deliverable evidentiarily irrelevant to the Texas one. The EU AI Act asks whether your training data is representative, and leans on the very geographic feature Illinois prohibits as a proxy.

Most of the audits are not there at all

Researchers who went looking for published Local Law 144 bias audits found them for 4.6% of 391 NYC employers, the finding they named Null Compliance (Cornell / Data & Society / Consumer Reports, FAccT 2024). Publishing the audit is the visible half of the obligation, and most of that sample had not reached it.

Enforcement is tightening around it

The NY State Comptroller found 17 potential Local Law 144 violations in the same 32-company sample where DCWP had found one, and DCWP agreed to shift to proactive enforcement (NY State Comptroller, December 2, 2025). The gap between a passing vendor audit and a defensible record is where that enforcement lands.

A bias audit is not the whole exposure

Scope self-classification, ADA accessibility on video interviews and FCRA adverse-action duties are separate legal theories a passed bias audit never tested. Mobley v. Workday, D.K. v. Intuit/HireVue and Kistler v. Eightfold each raise one of them, and none is decided.

The numbers and the verdicts are set in deterministic code, and the crew that narrates them cannot reach either.

The pipeline is short and the split of work inside it is strict. A vendor AEDT export is normalized into one canonical record schema, a deterministic engine computes every statistic, six regime rule packs and a policy gate decide every obligation, an agent crew writes the prose, and the run seals a hash-chained pre-audit package. Agents advise, code decides, which is what keeps the output filable regardless of how good or bad the model underneath is.

01 / THE DETERMINISTIC TRUST CORE

Every statistic computed outside the agent framework

A numpy engine computes marginal four-fifths impact ratios, intersectional race by sex ratios with Benjamini-Hochberg FDR control, protected-class proxy detection by Cramér's V, a counterfactual protected-attribute flip through a transparent logistic-regression surrogate, an ASR word-error-rate disparity on the video subset, and an FCRA trigger predicate. Same input, same output, every run.

02 / SIX RULE PACKS AND THE POLICY GATE

Four verdicts, decided in plain code

Each regime carries its own citation, effective date and required deliverable shape, and each obligation resolves to PASS, FAIL, NEEDS PROOF or CONFLICT. The gate is the one thing the language model cannot reach, so a well-argued narration can never move an obligation from NEEDS PROOF to green.

03 / THE AGENT CREW

Six narrators, a reconciler and a skeptic

Six jurisdiction agents narrate each regime's verdict, a conflict reconciler turns a hard conflict into a legal-strategy memo with the exposure of each branch, and an adversarial skeptic tries to refute every claimed pass and routes what it cannot clear to human proof. The crew is provider-swappable; with no provider configured it falls back to deterministic templates and the app runs identically.

04 / THE PRE-AUDIT PACKAGE

A bundle built to be signed by someone else

The run seals a SHA-256 hash-chained bundle of 17 nodes, each carrying its inputs, its computation and its rule citation plus a link to the hash of the node before it, so any edit breaks the chain. It exports as JSON and as a printable HTML auditor packet, and the chain is verified on every run and unit-tested for tamper-evidence.

The console is two views plus click-through dialogs rather than a single dashboard. Manual Testing holds the stage bar, the live Audit Execution trace and the candidate board. Run Benchmark, which stays disabled until Run Audit finishes, switches to Benchmark Results: three tiles reading Jurisdiction Deliverables 6, Lowest Impact Ratio 0.65 and Items Requiring Human Action 9, the six Jurisdiction Deliverables rows, a Supporting Evidence button row and Export Pre-Audit Package in the header.

Everything underneath opens as a dialog. Selection-Rate Analysis, the Conflict Register and the Human Proof Queue are the three Supporting Evidence buttons, and each regime's obligations and audit narrative open from its own deliverable row. Nothing is summarized into a single score, because a single score cannot be true across six differently shaped questions.

The vendor's audit passes. The audit the law actually asks for fails, on the same data.

The demo audits a seeded synthetic export of 1,040 candidates on requisition REQ-2026-0412, scored by three simulated vendor tools. The violation in it is planted and reproducible, built so the engine has something real to catch. Here is what the run puts on screen, in the order it puts it there.

The Clarion console with its Dataset Preview dialog open over the app bar, listing 16 representative records from what the dialog calls the deterministic 1,040-applicant audit population, which is seeded and synthetic. Each card shows a name, a role such as Operations Analyst or Warehouse Associate, race and sex cohort tags, a vendor score, and a green Advanced or red Screened Out decision. The header controls read About Demo, View Dataset, Run Audit and a greyed-out Run Benchmark.
What arrives before any audit runs. View Dataset opens 16 representative synthetic records, each a decision a vendor tool already made: a score and an advance or reject. No impact ratios appear here, because the app shows no vendor self-audit before the run. The whole audit works from this one normalized export, and Clarion never scores or ranks a candidate itself.
Clarion's Benchmark Results view. Three tiles read Jurisdiction Deliverables 6, Lowest Impact Ratio 0.65 and Items Requiring Human Action 9. Below a six-step Benchmark Execution trace, six Jurisdiction Deliverables rows read NYC Local Law 144 Live FAIL with 3 obligations, Colorado AI Act SB 24-205 effective Jun 30 2026 NEEDS PROOF with 2, Illinois HB 3773 Live CONFLICT with 2, Texas TRAIGA Live NEEDS PROOF with 2, California FEHA ADS amendments Live FAIL with 2, and EU AI Act Annex III effective Aug 2 2026 CONFLICT with 2. A Supporting Evidence row beneath offers Selection-Rate Analysis, Conflict Register and Human Proof Queue.
One audit, six shapes, no row of greens. Each row carries its own citation, effective date and deliverable format: an intersectional adverse-impact report for NYC, a reasonable-care impact assessment for Colorado, notice artifacts and proxy remediation for Illinois, an intent-based assessment for Texas, a records pack with a four-year retention attestation for California, and an Article 10 and Article 11 pack for the EU. Across those six deliverables sit 13 obligations: 2 PASS, 2 FAIL, 7 NEEDS PROOF and 2 CONFLICT.
Clarion's Selection-Rate Analysis dialog. An adverse-impact rate card lists White at a 51% advance rate and an impact ratio of 1.00, Asian at 49% and 0.95, Hispanic at 42% and 0.82, and Black at 42% and 0.82, with a line reading that the marginal race ratio 0.8196 and sex ratio 0.8744 both clear the 0.80 four-fifths line and the vendor stops here. A red FOUR-FIFTHS FAIL stamp beside it explains that cutting the same data by race and sex puts the Black / Female cohort at 0.65 against White Male. The Race × Sex Selection Grid (LL144) below shows White Male 1.00, White Female 0.96, Asian Male 0.96, Asian Female 0.91, Hispanic Male 0.84, Hispanic Female 0.76 in amber, Black Male 0.96, and Black Female 0.65 in red at 34% selected.
The turn. The marginal four-fifths test genuinely passes at 0.8196 for race and 0.8744 for sex. That is the audit a vendor self-report ships, and it really passes. Cut the same rows by race and sex, which is what Local Law 144 requires, and the Black / Female cell advances 44 of 130 at an impact ratio of 0.6471 against the White / Male reference cell. Under Benjamini-Hochberg FDR control at alpha 0.05 only that cell is statistically robust at q = 0.0185; Hispanic / Female at 0.7647 is flagged but not called significant, at q = 0.1629. The disparity is planted and synthetic, and the engine distinguishing a robust finding from a chance one is what a marginal-only audit has no way to produce.
Clarion's NYC Local Law 144 dialog, marked FAIL and Live, for the deliverable LL144 intersectional adverse-impact report plus posted public summary. Its Obligation Review lists marginal four-fifths impact ratios as PASS with race minimum 0.8196 and sex minimum 0.8744, intersectional race by sex impact ratios as FAIL with a Black / Female impact ratio of 0.6471 and failing cells Hispanic / Female and Black / Female, and a NEEDS PROOF item on scope self-classification requiring a human-counsel attestation. An audit narrative beneath explains the roll-up.
Every verdict opens into its obligations. The NYC row is one deliverable holding three obligations with three different statuses: the marginal ratios PASS at 0.8196 and 0.8744, the intersectional ratios FAIL at 0.6471, and the scope question sits at NEEDS PROOF until a named human answers it. The row keeps all three statuses side by side instead of rolling them into one, so a PASS on the marginal test never stands in for the intersectional test LL144 actually asks for.
Clarion's Conflict Register dialog reporting CONFLICT on zip_code and geographic data between il_hb3773 and eu_ai_act. The text explains that Illinois HB 3773 bans zip codes as protected-class proxies so the Illinois-safe configuration masks geography, while EU AI Act Article 10(3) requires relevant, representative and complete training data which typically requires geographic coverage, and that removing zip fails EU representativeness while keeping it violates Illinois. A recommendation proposes two configurations of one model, or knowingly accepting one exposure and documenting it, and states that the system will not certify fully compliant.
The conflict it will not bluff past. The engine measured zip_region, which the Conflict Register labels zip_code / geographic data, against race at a Cramér's V of 0.3321, over its 0.2 threshold, and cleared school_tier at 0.0711 in the same pass, so it is not flagging everything. Illinois bans the proxy; EU representativeness wants the same geography. One model configuration cannot satisfy both, so the gate emits CONFLICT and the reconciler writes the strategy memo instead: two deployment configurations, or one exposure accepted knowingly and documented, with EU high-risk penalties noted at their statutory maximum of the greater of 15 million euros or 3% of global annual turnover.
Clarion's Human Proof Queue dialog listing three skeptic challenges. AEDT scope for all scoring and filtering tools, against a vendor memo claiming the tool is not an AEDT, answered with self-classification is not a defense and routed to a human counsel scope attestation. A video interview ASR pipeline of the HireVue-style kind, against a claim that a passed race and sex bias audit covers it, answered with race and sex bias audit does not equal accessibility defense and routed to a human ADA and ASR accessibility review plus an accommodation workflow. A third-party-scored stream of the Eightfold-style kind, against a claim of no further exposure, answered with fairness of the score is irrelevant to FCRA and routed to FCRA adverse-action and dispute infrastructure.
Three refusals, each routed to a named human. On the 432 video-round candidates in the seeded synthetic export, the word error rate is 0.0794 for standard speech against 0.3016 for the 104 candidates with non-standard speech, a 3.8 times disparity that Local Law 144 never tests. Separately, 510 candidates were scored from third-party-scraped data and filtered on a numeric score, which trips the FCRA predicate: if the platform is a consumer reporting agency, every scored candidate is owed an adverse-action notice and a dispute path regardless of how fair the score is. Clarion detects and routes both. The accommodation workflow and the candidate-facing dispute portal named in those routes stay with the employer, and are not something we build here, and these three challenges are a different object from the 9 obligations counted as NEEDS PROOF or CONFLICT.
The exported Clarion pre-audit packet for Acme Logistics, Inc., listing jurisdictions NYC, Illinois and EU over 1040 candidates. Its integrity line reads SHA-256 hash chain, 17 nodes, with a tip hash and a green VERIFIED check. A coverage line reads 15% of 13 obligations auto-satisfied and 69% routed to human proof, with 2 fail, 7 needs-proof and 2 conflict, and one audit producing 6 jurisdiction deliverables. A note states that Veriprajna produces this pre-audit package and the independent LL144 sign-off remains with DCI Consulting, ORCAA or Secretariat. Per-regime verdict tables follow beneath.
The receipt. Export Pre-Audit Package emits the bundle as JSON and as this printable auditor packet: 17 hash-chained nodes with the chain verified on the run, every number carrying its computation and every verdict its citation. The coverage line reads 15% and 69%, which is 2 auto-satisfied obligations of 13 and the 9 that are NEEDS PROOF or CONFLICT of 13; the 2 FAIL obligations are in neither figure. The packet states in its own header that the independent sign-off is not ours.

What this demo claims, and what it deliberately does not.

Every figure on this page describes one seeded synthetic export of 1,040 candidates with a planted, reproducible violation. They demonstrate that the engine catches what a marginal-only audit misses. They are not an accuracy rate, not a benchmark against other products, and not a claim about any real employer's hiring data.

QuestionWhat Clarion does in this demoWhat remains outside the demo
CoverageOne audit run emits six jurisdiction-shaped deliverables across 13 obligations, with 2 PASS, 2 FAIL, 7 NEEDS PROOF and 2 CONFLICT.A compliance certificate or a fairness score. The gate is built to refuse one.
Adverse impactMarginal and intersectional four-fifths ratios with Benjamini-Hochberg FDR control, proxy detection by Cramér's V, and a counterfactual flip moving advance probability by 0.0574 on average and 0.1228 at maximum.Any statement about real hiring outcomes. The 1,040 candidates, the requisition and the disparity are synthetic and seeded.
Vendor dataReads one normalized AEDT export produced by Workday-Spotlight-style, HireVue-style and Eightfold-style scorers, served by stubbed fixture adapters.Live connectors into any ATS, assessment platform or match engine. No named vendor is a customer, partner or endorser.
Downstream actionDetects and routes the ADA and ASR accessibility finding and the FCRA adverse-action trigger to a named human, with the evidence attached.The CART accommodation workflow and the candidate-facing dispute portal. Both are simulated in the demo and are not something we build here.
Sign-offSeals a 17-node hash-chained pre-audit package, verified on the run and unit-tested for tamper-evidence, in JSON and printable HTML.The independent bias audit itself. Firms like DCI Consulting, ORCAA and Secretariat hold that role, and we do not.

What this demo does NOT do

Clarion does not certify compliance, and refusing to is the design. It is not a hiring model and never scores or ranks a candidate. The dataset is synthetic and seeded, the Black / Female disparity in it is planted so the engine has something real to catch rather than a finding about any actual employer, and the vendor connectors behind it are stubbed fixture adapters rather than live integrations. The tool produces evidence for counsel: it does not give legal advice or opine on liability, and Mobley v. Workday, D.K. v. Intuit/HireVue and Kistler v. Eightfold are pending theories rather than decided outcomes. Statutory maxima are quoted only as maxima, and no modeled exposure figure appears here. This page is an explainer with the walkthrough video, real screenshots, the mechanism and answers, not an app you operate from here.

What a General Counsel asks before putting a compliance layer over the hiring stack.

Our vendor already ran a bias audit and it passed. Why would we need a second one?

Because the test that passes and the test the law asks for are not the same test. On the demo's seeded 1,040-candidate synthetic export the marginal four-fifths check passes, with a minimum race impact ratio of 0.8196 and a minimum sex ratio of 0.8744 against the 0.80 line, and that is the audit a vendor self-report ships. NYC Local Law 144 asks for the intersectional race by sex ratios, and there the Black / Female cell advances 44 of 130 at an impact ratio of 0.6471. A marginal-only audit is not wrong about that cell, it is structurally blind to it.

Are you the auditor? Can you sign off on our Local Law 144 audit?

No. Veriprajna is not the independent auditor, and that role belongs to firms like DCI Consulting, ORCAA or Secretariat. Clarion produces the pre-audit package, a hash-chained bundle in which every number carries its computation and every verdict carries its citation, built so an independent auditor can sign it without rewriting it. The goal is to get the system into the state where the audit finds nothing worth writing up.

Does this connect to Workday or HireVue? We are not replacing our hiring stack.

It connects to neither, and Clarion never scores or ranks a candidate itself. Nothing in your stack is replaced. It reads a vendor AEDT export and normalizes it into one canonical record schema. In this demo the vendor connectors are stubbed fixture adapters over synthetic data, and the Workday-Spotlight-style scorer, the HireVue-style video round and the Eightfold-style match engine are archetypes rather than integrations.

There is an LLM in here. How do I put that in front of a regulator?

Every statistic, every threshold comparison, every conflict edge and the bundle hashing sit in deterministic code outside the agent framework. The six jurisdiction agents, the conflict reconciler and the adversarial skeptic write prose; they never compute a number and they cannot override the policy gate. With no model provider configured the crew falls back to deterministic templates and the app produces the same verdicts, because the verdicts were never the model's to give.

What happens when two regimes want opposite things? We need a decision, not a shrug.

You get the trade-off quantified and written down. In the demo the zip_region feature correlates with race at a Cramér's V of 0.3321, over the 0.2 threshold, so Illinois HB 3773 treats it as a banned proxy while EU AI Act Article 10(3) representativeness leans on the same geographic coverage. The gate emits CONFLICT instead of a pass, and the reconciler writes a strategy memo: run two deployment configurations, or accept one exposure knowingly and document it, noting that EU high-risk penalties reach a statutory maximum of the greater of 15 million euros or 3% of global annual turnover.

Nine of thirteen obligations need a human. Does this actually save us anything?

It automates the part that is honest to automate. On the demo's synthetic export 2 of 13 obligations are auto-satisfied, 2 fail outright, and 9 are routed as NEEDS PROOF or CONFLICT to a named human with the evidence already assembled and the citation attached. No honest system clears all thirteen on this data. A dashboard that showed thirteen greens would be clearing them without proof, and that dashboard is the exhibit a plaintiff reads back to you. The engine also clears school_tier as not a proxy at a Cramér's V of 0.0711, so it is not flagging everything it sees.

What do we actually hand the auditor at the end?

One export produces a pre-audit package as JSON and as a printable HTML auditor packet, version 1 of the veriprajna-pre-audit-package format, carrying 17 nodes with the chain verified on the run. Each node holds its inputs, its computation and its rule citation plus a SHA-256 link to the node before it, so any edit to any node breaks the chain. Alongside it sit the six jurisdiction deliverables, each shaped to its own regime's citation, effective date and required format.

Technical Research

The research behind this demo — the architecture, the verification design, and the enterprise blueprint.

Start with the two regimes your hiring stack may not be able to satisfy at the same time.

We are an AI engineering team, not a compliance certifier. We build the deterministic layer that computes the statistics, shapes the deliverable each regulator asked for, and names the contradictions out loud so counsel can decide with the exposure in front of them.

A useful first conversation is concrete: which AEDT vendors sit in your funnel, which of the six regimes you are actually exposed to, and which features in those models would look like protected-class proxies under one of them. We can work through the engine, the rule packs and the pre-audit package format alongside your HR, legal and data teams.

AEDT exposure assessment

  • ✓ Vendor tools that score or filter candidates
  • ✓ Which of the six regimes reach your funnel
  • ✓ Features that behave as protected-class proxies
  • ✓ Scope, ADA and FCRA theories a bias audit misses

Build the compliance overlay

  • ✓ Deterministic adverse-impact engine
  • ✓ Per-regime rule packs and the policy gate
  • ✓ Conflict reconciliation and human-proof routing
  • ✓ Hash-chained pre-audit package format