Project Red Team - reviewer audit pack

Hawthorne Arena Sample Audit

Release status

not release-safe

At least one check failed. The headline claims in this document should not be repeated in a public decision until the findings below are answered. This is a finding about what the document shows, not a finding that its numbers are wrong.

not release-safe: at least one check failed and the headline claims should not travel yet.

Audit reading: release status not release-safe; 4 claims inventoried; 5 evidence items not found; 0 findings contradicted by the document; 1 reading limit on the file; 13 evidence observations.

This pack and the board memo are rendered from one reading of one audit object. The line above appears in both and in the release proof; if the three ever differ, the pack is broken and not merely disputed.

Detector Status And Measured Limits

MeasureValue and limit
Detector modelearned model plus rules
Held-out recall76%
Representative-sample precision57% (about 4 in 10 flags were not what was sought)
Pooled precision82%, which is too flattering as a field rate
Observed document recall67% to 87%; intervals overlap and variation by document is not established
Lexicon limitPresence checks are word lists and word lists miss phrasings. Against a document we could check line by line, 3 of 12 presence checks returned a false negative. A negative result is 'we did not find it', never 'it is not there'.
Human reviewrequired before any finding here is treated as a conclusion about the study

Measured on the shipped audit predicate (rules/exemplar scorer plus the learned detector), on 3 held-out studies chosen by document hash: finds 76% of labelled qualifications. On the representative held-out sample it flagged 14 sentences and 6 were false positives, so about 4 in 10 flags were not qualifications; pooling the exhaustive positive-find pass with that sample gives 82% precision, which is useful but too flattering as a field-rate estimate. Observed document recalls ranged 67%-87%, but the intervals overlap; current evidence does not prove recall varies by house style. Cross-study comparison therefore needs the same extraction basis and a rank-stability check, not just raw counts.

A count in this pack may be set against another study's count only where three things hold: both were extracted on the same basis, the detector's measured error is printed beside both, and the ordering survives resampling. Where the ordering does not survive, we publish the band and the reason for it instead of a rank.

Source Status

audit-safe: text; plain-text extraction; 15 text lines; 0 tables; confidence 100% on the source-extraction scale

FieldValue
Source fileexamples/audit_pack/sample_study_excerpt.txt
Source kindtext
Source SHA-256f80071538816dcb78f3e8f718c487ec181425dedb7699bfaac6e8dfdf706f327
Extraction methodplain-text extraction
Source statusaudit-safe
Confidence100% on the source-extraction scale, where 100% is a document read without loss of text or tables
Pages0
Tables0
Text lines15
Table lines0
Notesnot recorded
Summaryaudit-safe: text; plain-text extraction; 15 text lines; 0 tables; confidence 100% on the source-extraction scale

Study Profile

FieldValue
TitleHawthorne Arena Sample Audit
Source fileexamples/audit_pack/sample_study_excerpt.txt
Study type as detectedmixed
Sector archetype as detectedtourism
GeographyHawthorne County, Oregon
Model named by the studyIMPLAN
Model yearnot recorded
Dollar yearnot recorded
Models detected in the textIMPLAN

Not Found In The Document As Submitted

Everything in this section is a miss by our detector, not an absence in the study. The detector reads for known phrasings and studies phrase things in ways no list anticipates. Read each line as "we did not find it", and close it by pointing us at the passage.

This section lists the 5 evidence items the audit read as a finding about the document. 1 further item is a limit on how much of the submitted file the audit could read, and is listed under Reading Limits On The Submitted File below. The register at the end of this pack lists all 6 items together, which is why its count is the larger one.

IDGateWeightFindingDetector's accountConsequenceWhat would close it
missing_public_cost_sidefiscal_truth_gateblockingNot found in the document as submitted: Public service costs, abatements, incentives, or other cost-side fiscal evidence.Tax or fiscal revenue appears without a detected public cost-side ledger.Gross tax revenue may be mistaken for net fiscal benefit.Point the audit at public service costs, abatements, incentives, or other cost-side fiscal evidence in the full report.
missing_survey_response_ratesurvey_adequacy_gateblockingNot found in the document as submitted: Survey response rate, sample frame, and usable response count.Survey language appears without response-rate or sample-size disclosure.Survey-derived spending or input assumptions cannot be weighted for coverage risk.Point the audit at survey response rate, sample frame, and usable response count in the full report.
missing_model_yearmodel_vintage_gateblockingNot found in the document as submitted: Model year or data year.IMPLAN is named, but no model/data year was detected.Model structure may not match the spending period.Point the audit at model year or data year in the full report.
missing_dollar_yearmodel_vintage_gateblockingNot found in the document as submitted: Dollar year or currency year.IMPLAN is named, but no dollar/currency year was detected.Nominal and real dollar claims may be mixed.Point the audit at dollar year or currency year in the full report.
missing_tax_retention_or_nettingtax_gross_vs_net_gateblockingNot found in the document as submitted: Gross-vs-net tax treatment and retention/abatement assumptions.Tax revenue appears without gross/net or tax-retention language.A gross tax collection can be mistaken for fiscal benefit to the decision maker.Point the audit at gross-vs-net tax treatment and retention/abatement assumptions in the full report.

Counted by detector: 76% recall on held-out studies, about 4 in 10 flags not what was sought, and per-document recall observed across 67% to 87% with overlapping intervals. Full limits are stated above.

Contradicted By Text The Document Contains

Everything in this section rests on text the study does contain. These are not misses: the document's own disclosure works against the claim attached to it, so they survive the detector's recall limit and a reader cannot close them by finding a passage we overlooked.

No finding of this kind was recorded.

Reading Limits On The Submitted File

This section is about our reading of the file you sent us, not about the study's contents. A thin extraction lowers what any finding below is worth, in both directions.

IDGateWeightFindingDetector's accountConsequenceWhat would close it
missing_reconcilable_tablessource_confidence_gatecautionNot found in the file as extracted: Tables or structured figures that can reconcile prose claims.No tables were detected by the existing document auditor.Prose figures cannot be checked against tabulated support.Send a cleaner file: a text-layer PDF, the original document, or the underlying tables.

Counted by detector: 76% recall on held-out studies, about 4 in 10 flags not what was sought, and per-document recall observed across 67% to 87% with overlapping intervals. Full limits are stated above.

Claim Inventory

Confidence is extraction confidence - how sure the audit is that it read the figure and its context correctly. It is not a judgment about whether the figure is right.

IDMetricValueExtraction confidenceSource clueSnippet
claim:001Jobs1,240 jobsmediumsame sentenceExecutive summary: the arena will create 1,240 jobs and generate $4.6 million in annual state and local tax revenue.
claim:002Tax and fiscal revenue$4.6 millionmediumsame sentenceExecutive summary: the arena will create 1,240 jobs and generate $4.6 million in annual state and local tax revenue.
claim:003Visitor spending$18.2 millionmediumsame sentenceThe arena attracts 310,000 visitors whose $18.2 million of spending is counted in the model.
claim:004Return on investment or benefit-cost ratio7.0:1 ratiomediumsame sentenceA $4.0 million city grant returns $28.0 million, a 7.0:1 ROI.

Counted by detector: 76% recall on held-out studies, about 4 in 10 flags not what was sought, and per-document recall observed across 67% to 87% with overlapping intervals. Full limits are stated above.

Gate Findings, Every Gate Run

GateStatusSubjectReasonConsequenceRelated
contribution_vs_impact_gatepassWhether the headline is new activity or activity that already existsDetected study type is mixed.Claim type is at least machine-labelable from the text.
fiscal_truth_gatefailWhether the tax figure is a net gain to the public purseFiscal revenue is detected, but no cost-side or net-fiscal treatment was detected.Do not release fiscal benefit language as net benefit.missing_public_cost_side
visitor_spending_gatepassWhether the visitor money is new to the area or already circulating in itVisitor/local or new-money language was detected.The split is addressed; a reviewer should still check the arithmetic.
survey_adequacy_gatefailWhether the survey behind the spending figure carries the weight put on itSurvey evidence is detected, but response coverage was not detected.Survey-based claims should carry a data-quality penalty.missing_survey_response_rate
model_vintage_gatefailWhether the study says which year's model and which year's dollars it usedThe model is named, but model year and/or dollar year is missing.Do not release year-sensitive claims without vintage disclosure.missing_model_year, missing_dollar_year
geography_fit_gatepassWhether the area modeled is the area the board is deciding forDetected study geography: Hawthorne County, Oregon.A geography is present for reviewer checks.
tax_gross_vs_net_gatefailWhether the tax figure is gross collections or money this jurisdiction keepsTax claims are detected, but gross-vs-net treatment was not detected.Tax claims should not be used as fiscal benefit until retention is shown.missing_tax_retention_or_netting
benefit_cost_gatepassWhether the return ratio has a defined cost sideBenefit-cost language includes cost, discount, and counterfactual cues.The ratio is reviewable from the text cues.
source_confidence_gatewarnHow much of the submitted document the audit was able to readSource cues exist, but no reconcilable tables were detected.Weigh every finding here against how much of the written standard could be checked mechanically on this document: 23 checks: 2 measured, 10 not matched, 11 need a reader.missing_reconcilable_tables

Counted by detector: 76% recall on held-out studies, about 4 in 10 flags not what was sought, and per-document recall observed across 67% to 87% with overlapping intervals. Full limits are stated above.

Every Expected-Evidence Item, As The Audit Object Records It

IDWeightExpectedWhy the audit says soConsequenceGate
missing_public_cost_sideblockingPublic service costs, abatements, incentives, or other cost-side fiscal evidenceTax or fiscal revenue appears without a detected public cost-side ledger.Gross tax revenue may be mistaken for net fiscal benefit.fiscal_truth_gate
missing_survey_response_rateblockingSurvey response rate, sample frame, and usable response countSurvey language appears without response-rate or sample-size disclosure.Survey-derived spending or input assumptions cannot be weighted for coverage risk.survey_adequacy_gate
missing_model_yearblockingModel year or data yearIMPLAN is named, but no model/data year was detected.Model structure may not match the spending period.model_vintage_gate
missing_dollar_yearblockingDollar year or currency yearIMPLAN is named, but no dollar/currency year was detected.Nominal and real dollar claims may be mixed.model_vintage_gate
missing_tax_retention_or_nettingblockingGross-vs-net tax treatment and retention/abatement assumptionsTax revenue appears without gross/net or tax-retention language.A gross tax collection can be mistaken for fiscal benefit to the decision maker.tax_gross_vs_net_gate
missing_reconcilable_tablescautionTables or structured figures that can reconcile prose claimsNo tables were detected by the existing document auditor.Prose figures cannot be checked against tabulated support.source_confidence_gate

Counted by detector: 76% recall on held-out studies, about 4 in 10 flags not what was sought, and per-document recall observed across 67% to 87% with overlapping intervals. Full limits are stated above.

Evidence Surface, Every Observation

Contract: evidence_surface.v1; audit contract: implan_audit_mode.v0.

ObservationKindMissingDirectionClaimConsequenceAuthorityProvenanceVintage
source:statusData quality penaltynoneutralSource extraction status: audit-safe.Sets the floor for what the audit may claim from this document; a weak source cannot support a strong release finding.Project Red Team source intakeexamples/audit_pack/sample_study_excerpt.txt
base:model_nameBase model inputnoneutralModel named by the study: IMPLAN.Sets the audit context that gates use before any headline claim can travel.IMPLAN Audit Mode study-profile extractorexamples/audit_pack/sample_study_excerpt.txt
base:geographyBase model inputnoneutralStudy geography named by the study: Hawthorne County, Oregon.Sets the audit context that gates use before any headline claim can travel.IMPLAN Audit Mode study-profile extractorexamples/audit_pack/sample_study_excerpt.txt
claim:001Hypothesis evidence atomnoneutralJobs claim: 1,240 jobs.Carried into the audit claim inventory; quote only with the attached snippet, source clue, confidence, and gate results.Extracted from report sentenceExecutive summary: the arena will create 1,240 jobs and generate $4.6 million in annual state and local tax revenue.
claim:002Hypothesis evidence atomnoneutralTax and fiscal revenue claim: $4.6 million.Carried into the audit claim inventory; quote only with the attached snippet, source clue, confidence, and gate results.Extracted from report sentenceExecutive summary: the arena will create 1,240 jobs and generate $4.6 million in annual state and local tax revenue.
claim:003Hypothesis evidence atomnoneutralVisitor spending claim: $18.2 million.Carried into the audit claim inventory; quote only with the attached snippet, source clue, confidence, and gate results.Extracted from report sentenceThe arena attracts 310,000 visitors whose $18.2 million of spending is counted in the model.
claim:004Hypothesis evidence atomnoneutralReturn on investment or benefit-cost ratio claim: 7.0:1 ratio.Carried into the audit claim inventory; quote only with the attached snippet, source clue, confidence, and gate results.Extracted from report sentenceA $4.0 million city grant returns $28.0 million, a 7.0:1 ROI.
missing_public_cost_sideMissing expected evidenceyesadverseMissing expected evidence: Public service costs, abatements, incentives, or other cost-side fiscal evidence.Gross tax revenue may be mistaken for net fiscal benefit.IMPLAN Audit Mode fiscal_truth_gate
missing_survey_response_rateMissing expected evidenceyesadverseMissing expected evidence: Survey response rate, sample frame, and usable response count.Survey-derived spending or input assumptions cannot be weighted for coverage risk.IMPLAN Audit Mode survey_adequacy_gate
missing_model_yearMissing expected evidenceyesadverseMissing expected evidence: Model year or data year.Model structure may not match the spending period.IMPLAN Audit Mode model_vintage_gate
missing_dollar_yearMissing expected evidenceyesadverseMissing expected evidence: Dollar year or currency year.Nominal and real dollar claims may be mixed.IMPLAN Audit Mode model_vintage_gate
missing_tax_retention_or_nettingMissing expected evidenceyesadverseMissing expected evidence: Gross-vs-net tax treatment and retention/abatement assumptions.A gross tax collection can be mistaken for fiscal benefit to the decision maker.IMPLAN Audit Mode tax_gross_vs_net_gate
missing_reconcilable_tablesMissing expected evidenceyesadverseMissing expected evidence: Tables or structured figures that can reconcile prose claims.Prose figures cannot be checked against tabulated support.IMPLAN Audit Mode source_confidence_gate

Counted by detector: 76% recall on held-out studies, about 4 in 10 flags not what was sought, and per-document recall observed across 67% to 87% with overlapping intervals. Full limits are stated above.

Gate Taxonomy

GatePurposeRequired EvidenceAliases
contribution_vs_impact_gateSeparate existing footprint, contribution, prospective impact, avoided loss, fiscal impact, and persuasion outcomes.claim type, time frame, baseline, new-to-geography status, counterfactual for prospective impactimpact vs contribution, new jobs, retained jobs, counterfactual, but-for, economic contribution
fiscal_truth_gateKeep gross tax output, net public revenue, public cost, incentive cost, and jurisdictional capture in separate lanes.tax level, recipient jurisdiction, period, public cost side, gross/net labelfiscal truth, public revenue, tax benefit, net fiscal impact, jurisdictional capture
visitor_spending_gateTest whether visitor spending is attributable, nonlocal, correctly margined, and not resident churn.visitor volume, spend per visitor, origin split, local/nonlocal split, displacement treatmentvisitor local split, nonlocal visitors, tourism spend, resident churn, retail margin
survey_adequacy_gateDowngrade claims resting on weak, heterogeneous, or non-representative surveys.sample frame, invited count, usable responses, response rate, weighting, collection datessurvey adequacy, sample size, response rate, proxy data, survey weighting
model_vintage_gateProtect trend and update claims from stale or incomparable model years, dollar years, geographies, and sector schemes.model year, source-data vintage, dollar year, deflator basis, sector scheme, geographymodel vintage, data year, dollar year, inflation basis, trend comparability
geography_fit_gateTest whether the model boundary matches the claim and whether sub-geography allocation is stated.boundary definition, geography IDs, allocation method, local purchase assumptions, spillover treatmentgeography fit, study area, decision geography, spillover, leakage
tax_gross_vs_net_gateCatch mislabeled tax arithmetic wherever it appears.tax base, rate or model source, recipient, period, construction/operations split, gross or net labelgross vs net tax, tax retention, abatement, incremental fiscal, recipient jurisdiction
benefit_cost_gatePrevent ROI, BCR, payback, and public-return ratios from mixing economic activity with public benefit or omitting public cost.numerator, denominator, payer, beneficiary, public cost, counterfactual, discount rateBCA, ROI, benefit-cost ratio, payback, public return
source_confidence_gateSet a claim floor from source type, provenance, vintage, and extractability.source owner, vintage, locator, extraction method, checksum when known, attestation for client-supplied inputssource confidence, OCR confidence, table recovery, checksum, client-supplied inputs

Artifact Manifest And Hashes

FieldValue
Source fileexamples/audit_pack/sample_study_excerpt.txt
Source SHA-256f80071538816dcb78f3e8f718c487ec181425dedb7699bfaac6e8dfdf706f327
Extraction methodplain-text extraction
Audit contractimplan_audit_mode.v0
Evidence surface contractevidence_surface.v1

Per-artifact hashes are not embedded in this file: the manifest hashes this page, so it cannot contain its own digest. They are written to manifest.json beside this pack, one SHA-256 per artifact. A digest that no longer matches means the file was edited after the run, and the pack should not be treated as evidence of what the audit produced.

Negative results are detector results: not matched means not found at the current extraction and recall limits, not proof that the submitted study contains no such evidence. Only the contradicted section rests on text the document does contain.