Project Red Team - reviewer audit pack
Hawthorne Arena Sample Audit
Release status
not release-safe
At least one check failed. The headline claims in this document should not be repeated in a public decision until the findings below are answered. This is a finding about what the document shows, not a finding that its numbers are wrong.
not release-safe: at least one check failed and the headline claims should not travel yet.
Audit reading: release status not release-safe; 4 claims inventoried; 5 evidence items not found; 0 findings contradicted by the document; 1 reading limit on the file; 13 evidence observations.
This pack and the board memo are rendered from one reading of one audit object. The line above appears in both and in the release proof; if the three ever differ, the pack is broken and not merely disputed.
Detector Status And Measured Limits
| Measure | Value and limit |
|---|---|
| Detector mode | learned model plus rules |
| Held-out recall | 76% |
| Representative-sample precision | 57% (about 4 in 10 flags were not what was sought) |
| Pooled precision | 82%, which is too flattering as a field rate |
| Observed document recall | 67% to 87%; intervals overlap and variation by document is not established |
| Lexicon limit | Presence checks are word lists and word lists miss phrasings. Against a document we could check line by line, 3 of 12 presence checks returned a false negative. A negative result is 'we did not find it', never 'it is not there'. |
| Human review | required before any finding here is treated as a conclusion about the study |
Measured on the shipped audit predicate (rules/exemplar scorer plus the learned detector), on 3 held-out studies chosen by document hash: finds 76% of labelled qualifications. On the representative held-out sample it flagged 14 sentences and 6 were false positives, so about 4 in 10 flags were not qualifications; pooling the exhaustive positive-find pass with that sample gives 82% precision, which is useful but too flattering as a field-rate estimate. Observed document recalls ranged 67%-87%, but the intervals overlap; current evidence does not prove recall varies by house style. Cross-study comparison therefore needs the same extraction basis and a rank-stability check, not just raw counts.
A count in this pack may be set against another study's count only where three things hold: both were extracted on the same basis, the detector's measured error is printed beside both, and the ordering survives resampling. Where the ordering does not survive, we publish the band and the reason for it instead of a rank.
Source Status
audit-safe: text; plain-text extraction; 15 text lines; 0 tables; confidence 100% on the source-extraction scale
| Field | Value |
|---|---|
| Source file | examples/audit_pack/sample_study_excerpt.txt |
| Source kind | text |
| Source SHA-256 | f80071538816dcb78f3e8f718c487ec181425dedb7699bfaac6e8dfdf706f327 |
| Extraction method | plain-text extraction |
| Source status | audit-safe |
| Confidence | 100% on the source-extraction scale, where 100% is a document read without loss of text or tables |
| Pages | 0 |
| Tables | 0 |
| Text lines | 15 |
| Table lines | 0 |
| Notes | not recorded |
| Summary | audit-safe: text; plain-text extraction; 15 text lines; 0 tables; confidence 100% on the source-extraction scale |
Study Profile
| Field | Value |
|---|---|
| Title | Hawthorne Arena Sample Audit |
| Source file | examples/audit_pack/sample_study_excerpt.txt |
| Study type as detected | mixed |
| Sector archetype as detected | tourism |
| Geography | Hawthorne County, Oregon |
| Model named by the study | IMPLAN |
| Model year | not recorded |
| Dollar year | not recorded |
| Models detected in the text | IMPLAN |
Not Found In The Document As Submitted
Everything in this section is a miss by our detector, not an absence in the study. The detector reads for known phrasings and studies phrase things in ways no list anticipates. Read each line as "we did not find it", and close it by pointing us at the passage.
This section lists the 5 evidence items the audit read as a finding about the document. 1 further item is a limit on how much of the submitted file the audit could read, and is listed under Reading Limits On The Submitted File below. The register at the end of this pack lists all 6 items together, which is why its count is the larger one.
| ID | Gate | Weight | Finding | Detector's account | Consequence | What would close it |
|---|---|---|---|---|---|---|
| missing_public_cost_side | fiscal_truth_gate | blocking | Not found in the document as submitted: Public service costs, abatements, incentives, or other cost-side fiscal evidence. | Tax or fiscal revenue appears without a detected public cost-side ledger. | Gross tax revenue may be mistaken for net fiscal benefit. | Point the audit at public service costs, abatements, incentives, or other cost-side fiscal evidence in the full report. |
| missing_survey_response_rate | survey_adequacy_gate | blocking | Not found in the document as submitted: Survey response rate, sample frame, and usable response count. | Survey language appears without response-rate or sample-size disclosure. | Survey-derived spending or input assumptions cannot be weighted for coverage risk. | Point the audit at survey response rate, sample frame, and usable response count in the full report. |
| missing_model_year | model_vintage_gate | blocking | Not found in the document as submitted: Model year or data year. | IMPLAN is named, but no model/data year was detected. | Model structure may not match the spending period. | Point the audit at model year or data year in the full report. |
| missing_dollar_year | model_vintage_gate | blocking | Not found in the document as submitted: Dollar year or currency year. | IMPLAN is named, but no dollar/currency year was detected. | Nominal and real dollar claims may be mixed. | Point the audit at dollar year or currency year in the full report. |
| missing_tax_retention_or_netting | tax_gross_vs_net_gate | blocking | Not found in the document as submitted: Gross-vs-net tax treatment and retention/abatement assumptions. | Tax revenue appears without gross/net or tax-retention language. | A gross tax collection can be mistaken for fiscal benefit to the decision maker. | Point the audit at gross-vs-net tax treatment and retention/abatement assumptions in the full report. |
Counted by detector: 76% recall on held-out studies, about 4 in 10 flags not what was sought, and per-document recall observed across 67% to 87% with overlapping intervals. Full limits are stated above.
Contradicted By Text The Document Contains
Everything in this section rests on text the study does contain. These are not misses: the document's own disclosure works against the claim attached to it, so they survive the detector's recall limit and a reader cannot close them by finding a passage we overlooked.
No finding of this kind was recorded.
Reading Limits On The Submitted File
This section is about our reading of the file you sent us, not about the study's contents. A thin extraction lowers what any finding below is worth, in both directions.
| ID | Gate | Weight | Finding | Detector's account | Consequence | What would close it |
|---|---|---|---|---|---|---|
| missing_reconcilable_tables | source_confidence_gate | caution | Not found in the file as extracted: Tables or structured figures that can reconcile prose claims. | No tables were detected by the existing document auditor. | Prose figures cannot be checked against tabulated support. | Send a cleaner file: a text-layer PDF, the original document, or the underlying tables. |
Counted by detector: 76% recall on held-out studies, about 4 in 10 flags not what was sought, and per-document recall observed across 67% to 87% with overlapping intervals. Full limits are stated above.
Claim Inventory
Confidence is extraction confidence - how sure the audit is that it read the figure and its context correctly. It is not a judgment about whether the figure is right.
| ID | Metric | Value | Extraction confidence | Source clue | Snippet |
|---|---|---|---|---|---|
| claim:001 | Jobs | 1,240 jobs | medium | same sentence | Executive summary: the arena will create 1,240 jobs and generate $4.6 million in annual state and local tax revenue. |
| claim:002 | Tax and fiscal revenue | $4.6 million | medium | same sentence | Executive summary: the arena will create 1,240 jobs and generate $4.6 million in annual state and local tax revenue. |
| claim:003 | Visitor spending | $18.2 million | medium | same sentence | The arena attracts 310,000 visitors whose $18.2 million of spending is counted in the model. |
| claim:004 | Return on investment or benefit-cost ratio | 7.0:1 ratio | medium | same sentence | A $4.0 million city grant returns $28.0 million, a 7.0:1 ROI. |
Counted by detector: 76% recall on held-out studies, about 4 in 10 flags not what was sought, and per-document recall observed across 67% to 87% with overlapping intervals. Full limits are stated above.
Gate Findings, Every Gate Run
| Gate | Status | Subject | Reason | Consequence | Related |
|---|---|---|---|---|---|
| contribution_vs_impact_gate | pass | Whether the headline is new activity or activity that already exists | Detected study type is mixed. | Claim type is at least machine-labelable from the text. | — |
| fiscal_truth_gate | fail | Whether the tax figure is a net gain to the public purse | Fiscal revenue is detected, but no cost-side or net-fiscal treatment was detected. | Do not release fiscal benefit language as net benefit. | missing_public_cost_side |
| visitor_spending_gate | pass | Whether the visitor money is new to the area or already circulating in it | Visitor/local or new-money language was detected. | The split is addressed; a reviewer should still check the arithmetic. | — |
| survey_adequacy_gate | fail | Whether the survey behind the spending figure carries the weight put on it | Survey evidence is detected, but response coverage was not detected. | Survey-based claims should carry a data-quality penalty. | missing_survey_response_rate |
| model_vintage_gate | fail | Whether the study says which year's model and which year's dollars it used | The model is named, but model year and/or dollar year is missing. | Do not release year-sensitive claims without vintage disclosure. | missing_model_year, missing_dollar_year |
| geography_fit_gate | pass | Whether the area modeled is the area the board is deciding for | Detected study geography: Hawthorne County, Oregon. | A geography is present for reviewer checks. | — |
| tax_gross_vs_net_gate | fail | Whether the tax figure is gross collections or money this jurisdiction keeps | Tax claims are detected, but gross-vs-net treatment was not detected. | Tax claims should not be used as fiscal benefit until retention is shown. | missing_tax_retention_or_netting |
| benefit_cost_gate | pass | Whether the return ratio has a defined cost side | Benefit-cost language includes cost, discount, and counterfactual cues. | The ratio is reviewable from the text cues. | — |
| source_confidence_gate | warn | How much of the submitted document the audit was able to read | Source cues exist, but no reconcilable tables were detected. | Weigh every finding here against how much of the written standard could be checked mechanically on this document: 23 checks: 2 measured, 10 not matched, 11 need a reader. | missing_reconcilable_tables |
Counted by detector: 76% recall on held-out studies, about 4 in 10 flags not what was sought, and per-document recall observed across 67% to 87% with overlapping intervals. Full limits are stated above.
Every Expected-Evidence Item, As The Audit Object Records It
| ID | Weight | Expected | Why the audit says so | Consequence | Gate |
|---|---|---|---|---|---|
| missing_public_cost_side | blocking | Public service costs, abatements, incentives, or other cost-side fiscal evidence | Tax or fiscal revenue appears without a detected public cost-side ledger. | Gross tax revenue may be mistaken for net fiscal benefit. | fiscal_truth_gate |
| missing_survey_response_rate | blocking | Survey response rate, sample frame, and usable response count | Survey language appears without response-rate or sample-size disclosure. | Survey-derived spending or input assumptions cannot be weighted for coverage risk. | survey_adequacy_gate |
| missing_model_year | blocking | Model year or data year | IMPLAN is named, but no model/data year was detected. | Model structure may not match the spending period. | model_vintage_gate |
| missing_dollar_year | blocking | Dollar year or currency year | IMPLAN is named, but no dollar/currency year was detected. | Nominal and real dollar claims may be mixed. | model_vintage_gate |
| missing_tax_retention_or_netting | blocking | Gross-vs-net tax treatment and retention/abatement assumptions | Tax revenue appears without gross/net or tax-retention language. | A gross tax collection can be mistaken for fiscal benefit to the decision maker. | tax_gross_vs_net_gate |
| missing_reconcilable_tables | caution | Tables or structured figures that can reconcile prose claims | No tables were detected by the existing document auditor. | Prose figures cannot be checked against tabulated support. | source_confidence_gate |
Counted by detector: 76% recall on held-out studies, about 4 in 10 flags not what was sought, and per-document recall observed across 67% to 87% with overlapping intervals. Full limits are stated above.
Evidence Surface, Every Observation
Contract: evidence_surface.v1; audit contract: implan_audit_mode.v0.
| Observation | Kind | Missing | Direction | Claim | Consequence | Authority | Provenance | Vintage |
|---|---|---|---|---|---|---|---|---|
| source:status | Data quality penalty | no | neutral | Source extraction status: audit-safe. | Sets the floor for what the audit may claim from this document; a weak source cannot support a strong release finding. | Project Red Team source intake | examples/audit_pack/sample_study_excerpt.txt | — |
| base:model_name | Base model input | no | neutral | Model named by the study: IMPLAN. | Sets the audit context that gates use before any headline claim can travel. | IMPLAN Audit Mode study-profile extractor | examples/audit_pack/sample_study_excerpt.txt | — |
| base:geography | Base model input | no | neutral | Study geography named by the study: Hawthorne County, Oregon. | Sets the audit context that gates use before any headline claim can travel. | IMPLAN Audit Mode study-profile extractor | examples/audit_pack/sample_study_excerpt.txt | — |
| claim:001 | Hypothesis evidence atom | no | neutral | Jobs claim: 1,240 jobs. | Carried into the audit claim inventory; quote only with the attached snippet, source clue, confidence, and gate results. | Extracted from report sentence | Executive summary: the arena will create 1,240 jobs and generate $4.6 million in annual state and local tax revenue. | — |
| claim:002 | Hypothesis evidence atom | no | neutral | Tax and fiscal revenue claim: $4.6 million. | Carried into the audit claim inventory; quote only with the attached snippet, source clue, confidence, and gate results. | Extracted from report sentence | Executive summary: the arena will create 1,240 jobs and generate $4.6 million in annual state and local tax revenue. | — |
| claim:003 | Hypothesis evidence atom | no | neutral | Visitor spending claim: $18.2 million. | Carried into the audit claim inventory; quote only with the attached snippet, source clue, confidence, and gate results. | Extracted from report sentence | The arena attracts 310,000 visitors whose $18.2 million of spending is counted in the model. | — |
| claim:004 | Hypothesis evidence atom | no | neutral | Return on investment or benefit-cost ratio claim: 7.0:1 ratio. | Carried into the audit claim inventory; quote only with the attached snippet, source clue, confidence, and gate results. | Extracted from report sentence | A $4.0 million city grant returns $28.0 million, a 7.0:1 ROI. | — |
| missing_public_cost_side | Missing expected evidence | yes | adverse | Missing expected evidence: Public service costs, abatements, incentives, or other cost-side fiscal evidence. | Gross tax revenue may be mistaken for net fiscal benefit. | IMPLAN Audit Mode fiscal_truth_gate | — | — |
| missing_survey_response_rate | Missing expected evidence | yes | adverse | Missing expected evidence: Survey response rate, sample frame, and usable response count. | Survey-derived spending or input assumptions cannot be weighted for coverage risk. | IMPLAN Audit Mode survey_adequacy_gate | — | — |
| missing_model_year | Missing expected evidence | yes | adverse | Missing expected evidence: Model year or data year. | Model structure may not match the spending period. | IMPLAN Audit Mode model_vintage_gate | — | — |
| missing_dollar_year | Missing expected evidence | yes | adverse | Missing expected evidence: Dollar year or currency year. | Nominal and real dollar claims may be mixed. | IMPLAN Audit Mode model_vintage_gate | — | — |
| missing_tax_retention_or_netting | Missing expected evidence | yes | adverse | Missing expected evidence: Gross-vs-net tax treatment and retention/abatement assumptions. | A gross tax collection can be mistaken for fiscal benefit to the decision maker. | IMPLAN Audit Mode tax_gross_vs_net_gate | — | — |
| missing_reconcilable_tables | Missing expected evidence | yes | adverse | Missing expected evidence: Tables or structured figures that can reconcile prose claims. | Prose figures cannot be checked against tabulated support. | IMPLAN Audit Mode source_confidence_gate | — | — |
Counted by detector: 76% recall on held-out studies, about 4 in 10 flags not what was sought, and per-document recall observed across 67% to 87% with overlapping intervals. Full limits are stated above.
Gate Taxonomy
| Gate | Purpose | Required Evidence | Aliases |
|---|---|---|---|
| contribution_vs_impact_gate | Separate existing footprint, contribution, prospective impact, avoided loss, fiscal impact, and persuasion outcomes. | claim type, time frame, baseline, new-to-geography status, counterfactual for prospective impact | impact vs contribution, new jobs, retained jobs, counterfactual, but-for, economic contribution |
| fiscal_truth_gate | Keep gross tax output, net public revenue, public cost, incentive cost, and jurisdictional capture in separate lanes. | tax level, recipient jurisdiction, period, public cost side, gross/net label | fiscal truth, public revenue, tax benefit, net fiscal impact, jurisdictional capture |
| visitor_spending_gate | Test whether visitor spending is attributable, nonlocal, correctly margined, and not resident churn. | visitor volume, spend per visitor, origin split, local/nonlocal split, displacement treatment | visitor local split, nonlocal visitors, tourism spend, resident churn, retail margin |
| survey_adequacy_gate | Downgrade claims resting on weak, heterogeneous, or non-representative surveys. | sample frame, invited count, usable responses, response rate, weighting, collection dates | survey adequacy, sample size, response rate, proxy data, survey weighting |
| model_vintage_gate | Protect trend and update claims from stale or incomparable model years, dollar years, geographies, and sector schemes. | model year, source-data vintage, dollar year, deflator basis, sector scheme, geography | model vintage, data year, dollar year, inflation basis, trend comparability |
| geography_fit_gate | Test whether the model boundary matches the claim and whether sub-geography allocation is stated. | boundary definition, geography IDs, allocation method, local purchase assumptions, spillover treatment | geography fit, study area, decision geography, spillover, leakage |
| tax_gross_vs_net_gate | Catch mislabeled tax arithmetic wherever it appears. | tax base, rate or model source, recipient, period, construction/operations split, gross or net label | gross vs net tax, tax retention, abatement, incremental fiscal, recipient jurisdiction |
| benefit_cost_gate | Prevent ROI, BCR, payback, and public-return ratios from mixing economic activity with public benefit or omitting public cost. | numerator, denominator, payer, beneficiary, public cost, counterfactual, discount rate | BCA, ROI, benefit-cost ratio, payback, public return |
| source_confidence_gate | Set a claim floor from source type, provenance, vintage, and extractability. | source owner, vintage, locator, extraction method, checksum when known, attestation for client-supplied inputs | source confidence, OCR confidence, table recovery, checksum, client-supplied inputs |
Artifact Manifest And Hashes
| Field | Value |
|---|---|
| Source file | examples/audit_pack/sample_study_excerpt.txt |
| Source SHA-256 | f80071538816dcb78f3e8f718c487ec181425dedb7699bfaac6e8dfdf706f327 |
| Extraction method | plain-text extraction |
| Audit contract | implan_audit_mode.v0 |
| Evidence surface contract | evidence_surface.v1 |
Per-artifact hashes are not embedded in this file: the manifest hashes this page, so it cannot contain its own digest. They are written to manifest.json beside this pack, one SHA-256 per artifact. A digest that no longer matches means the file was edited after the run, and the pack should not be treated as evidence of what the audit produced.
Negative results are detector results: not matched means not found at the current extraction and recall limits, not proof that the submitted study contains no such evidence. Only the contradicted section rests on text the document does contain.