# Million-scale foundation and real document run Date: `2026-08-10` Updated: `2026-08-13` > Integrity correction (`2026-08-12`): this report preserves historical engineering context, but every statement below based on synthetic/mock/fallback data is `RETIRED_NON_EVIDENCE` and must not be cited as current coverage, capacity, readiness, privacy, or publication proof. Those data artifacts were removed. Current authority is `ROADMAP.md`, `docs/etl/e2e-scrape-load-tracker.md`, `docs/etl/real-corpus-registry.json`, and `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/non-real-record-purge-20260812.json`. The clean ledger is `126,760` real rows in `13` Parquet files. Official public-domain identities are retained exactly; v3 identity-suppression claims are retired. ## Where we are now - The synthetic durable-queue benchmark completes `1,000,000/1,000,000` work items with bounded memory and zero unfinished work. - The Docker ETL build context is now bounded independently of local corpus, evidence archives, package caches, and sparse root databases. Earlier runs exposed `4.55 GB`, `87.80 MB`, and `514.20 MB` regressions. The final regression run then exposed repo-local `.cache`, `.local`, `tmp`, `outputs`, and root `*.db` leakage at `706.08 MB`. `.dockerignore` now excludes those generated paths while retaining the official raw samples and the required published election fixture; the next Docker build transferred `709.70 kB`, and the final incremental regression build transferred `368.25 kB`. The full container suite completed `1,303` tests with `11` real-capture-dependent skips and zero failures. No corpus or unrelated Docker artifact was deleted. - The document lane now has a separate million-manifest capacity proof using realistic synthetic PDF metadata across `32` host partitions. A `1,088,696,320`-byte SQLite queue ingests `1,000,000` unique manifests at `45,258.633/s`; exact replay inserts `0` at `42,242.245/s`; indexed claim/complete terminates every item at `6,424.421/s`; JSON, distinct-key, queue, SQLite, and FK reconciliation all pass with `53.359 MB` process peak RSS. A direct streaming-CAS subgate transfers `128 MiB` in `64 KiB` chunks, materializes `64` unique SHA-256 objects, deduplicates all `64` repeats, rejects oversize before transfer, and leaves zero partials. It makes zero network requests and explicitly denies real document coverage, real network behavior, and real extraction/OCR quality. - The production document-fetch worker now passes a controlled loopback HTTP matrix. All `10` primary work items terminate as `5` succeeded / `5` dead across `11` queue attempts: `503` recovers on request three, `429` on request two with `Retry-After`, a deliberately truncated body recovers on request two, `404` and declared oversize fail closed, and two `403` responses open a cohort circuit that prevents the third blocked request. An isolated socket-timeout probe exceeds `1s` on the first response and succeeds on the second internal download attempt while keeping one queue attempt, one CAS object, and zero partials. Five linked primary documents reconcile to five source records/text rows and four CAS objects because the shared payload deduplicates; SQLite `quick_check` is `ok`, FK errors are zero, and no partial files remain. A concurrent heartbeat subgate keeps a `4.5s` download under its original `3s` lease: a competing claim at `3.4s` receives zero items, and the owner finishes succeeded in one attempt. The initial matrix run exposed that a body shorter than its declared `Content-Length` was accepted. The streaming store now rejects that transfer atomically, deletes the partial, and lets the queue retry it. This matrix makes zero external requests and proves no real-document coverage. - Link discovery is separated from network fetch. `4,035` initiatives were scanned and `9,084` document URL rows were materialized without downloading bytes. - The real Senate document run fetched `6,792/9,084` distinct documents, classified `198` permanent failures, and retained `2,094` retryable items. - Text extraction completed `6,792/6,792` fetched documents: `4,740` XML and `2,052` HTML, totaling `29,728,661` normalized text characters with zero truncations. - All `3,607` Senate initiatives that expose document links have at least one fetched document. - Local content-addressed replication contains `6,792` objects / `133,219,457` bytes. A deterministic `100`-object restore sample passed checksum verification. - The broader local real-document inventory now discovers `111,422` files and classifies `111,408` accountability-eligible official instances / `109,558` distinct content objects (`5,093,218,479` bytes) after excluding `5` fixture/sample files and `9` preserved real transport-failure captures. BOE now contains `1,034` exact daily summaries and `89,001` individual XML documents. The historical candidate expansion added `18` exact summary dates and `21` exact documents; every document job is terminal success and every BOE document retains summary item, HTTPS URL, bytes, SHA-256, retrieval metadata, and portable manifest. Incremental inventory reuses `111,369` checksum-identical rows and inspects only `48`. All `1,437` PDFs parse across `44,927` pages. Provenance balances all `111,408` eligible rows, verifies checksum and public URL for `111,357` (`99.9542%`), and reports zero conflicts. The same `51` historical files remain unlinked and unpromoted; no BOE document added a gap. - The complete current real-document corpus runs through the disk-backed resumable v5 native-text pipeline. It verifies every raw SHA-256, deduplicates by content, leases bounded claims, writes deterministic gzip CAS atomically, retains unsanitized official public identity, and manifests every file instance. The run terminates `109,558/109,558` distinct contents and `111,408/111,408` instances with `1,999,902,681` native characters, `108,139` unique text CAS objects / `533,125,716` compressed bytes, zero retries, dead work, failures, or truncation. An independent validator rehashes every raw instance and CAS object. - Additive OCR now processes the entire routed current corpus rather than a sample: `856/856` empty or sparse pages across `58` real PDFs, using persisted input-manifest SHA, engine/language/version/configuration, durable work state, and separate OCR CAS. It produces text on `849` pages, keeps `7` explicit no-text results, emits `89,632` characters in `131` unique CAS objects, and records `17` pages with at least a 50-character gain. An independent validator rehashes all `58` source PDFs and all `131` OCR objects and proves native text was not replaced. A deterministic `40`-page human packet selects all seven no-text results, all 17 material gains, and both source/routing groups; decision, note, reviewer, and date fields remain blank. Human quality is honestly `0/40`, so semantic OCR accuracy and analytical promotion remain open. - The vote-source transport audit validates all `1,809,222` rows and `8,373` shard checksums, then classifies `6,426` distinct official source URLs without rewriting historical evidence. HTTPS covers `1,707,050` rows / `5,260` URLs. The remaining `102,172` HTTP rows / `1,166` URLs are confined to Senate legislatures 10 and 12: `33,683` rows / `484` URLs have local checksum captures and `68,489` rows / `682` URLs lack an immutable replacement. Two explicit HTTPS-equivalence probes returned `403` HTML (`417/418` bytes), so they prove no equivalence and remain an open public-access incident. Routine audit is offline via `just etl-scale-audit-vote-source-urls`; another network probe requires a new lever. - The real inventory passes the numeric `S1 100,000` gate at `111,408`. Native processing and OCR routing are complete and independently validated. Representative scanned/office/language/host/size coverage, `40/40` human OCR quality, cost per `1,000`, full remote raw-object origin, and one-million documents remain open. - The exact BOE candidate base now spans seven immutable proclamation documents, `4,901` candidacies and `42,056` occurrences: `25,077` titular and `16,979` substitute rows. Eleven correction documents retain all `956` exact paragraphs. The reviewed three-proclamation 2023/2024 subset remains `10,815` base rows / `1,188` candidacies; three applied corrections retain `180` paragraphs, of which `127` support `17` operations and `53` remain structural/context. Effective state is `10,820` normalized rows with complete person/party/source/document lineage, `10,795` exact base person links, `25` new/changed source-scoped identities, and `88` immutable correction-backed source records. Eight correction documents / `776` paragraphs remain `captured_unapplied`; their four proclamation documents / `3,713` candidacies / `31,241` candidates are barred from effective state. November 2019 is captured but candidate parsing fails closed where the official heading lacks its required candidacy number. The v3 semantic artifact remains two Zstandard Parquet files (`2,083,541` bytes), validates and replays `2/2`, and keeps `is_elected=null`. Every base public name and paragraph remains intact; no missing identifier, cross-source identity, or outcome is inferred. This is real local evidence, not representative scope, durable origin, clean restore, or one-million scale. - The pending historical correction lane is now executable instead of narrative. A deterministic CSV contains `776/776` exact official paragraphs as immutable dual-review tasks with blank decisions, URL/checksum/text lineage, stable fingerprints and structural candidate hints. A separate deduplicated JSONL retains the full `3,713` candidacies and `31,241` public candidate occurrences once (`14,686,321` bytes) instead of duplicating candidate lists per task; the review CSV is `1,186,277` bytes. Regeneration preserves completed decisions only when the task fingerprint is unchanged and fails closed on evidence/schema removal or drift. Central readiness independently rehashes both artifacts, validates every public name and source text, and reconciles them against the pending database rows. No effective row or correction operation was changed. - Five old-legislature detail-route circuits opened after repeated `HTTP 500`; `2,070` later requests were skipped locally. No further network retry is allowed in this sprint without a new lever. - The legacy million-vote audit now treats Senate `AUSENTE` as absence, not `no_vote`: `5,437/6,947` comparable snapshot events reconcile (`78.2640%`), up from the semantically invalid prior result. The legacy snapshot still lacks explicit absence totals and remains unpromoted. - A clean offline repair on a disposable database isolates cache lookup by legislature, blocks HTTP at the loader, preserves `totals_absent`, and removes stale stored seats after successful authoritative parses. It processed `4,313` cached Senate events (`4,205` with member votes), loaded/upserted `1,113,152` member rows, removed `5,319` stale rows, skipped `1,221` uncached events, and retained `17` no-match events with zero fetch failures. - The repaired database contains `1,877,033` member votes, passes SQLite integrity with `0` FK errors (baseline `44,646`), and reconciles `5,830/6,952` comparable events (`83.8608%`). Audit v2 classifies the residual: `1,115/1,122` mismatches contain fewer observed category rows than the comparable official totals, `7/1,122` preserve the comparable category population but distribute it differently, and none contain excess rows; `1,111` mismatches are Senate and `11` Congress. The exact-label identity bridge was changed from one full update scan per label to one bounded unresolved-row scan plus indexed temp-map lookups; it links `100%` of member rows and adds `948` observed-label people / `3,360` net observed mandates. It is proof, not promotion: exact labels are not verified external identities, the database is disposable/unpublished, and the missing rows still require authoritative capture. - A new SQLite-direct publisher avoids the monolithic snapshot entirely. It emitted `8,357` deterministic event shards containing `1,877,033` member votes in `179,993,722` gzip bytes, with a `4,528,365`-byte index and a maximum shard of `31,761` bytes / `350` member votes. Independent validation passed every file, checksum, payload, total, source-record lineage, public URL, and private-path check. Publication remains local-only. - A typed semantic Parquet lane now streams all `1,877,033` member votes into `20` snapshot/source/jurisdiction/year partitions and `33` Zstandard files (`10,795,673` bytes). Export peak RSS is `122.625 MB`; independent full-row/hash/schema/privacy validation peak RSS is `226.078 MB`. Source-record lineage and public URL coverage are `100%`, with URL scope kept explicit: `732,975` record-specific and `1,144,058` official source-default rows. An unchanged next-snapshot build reused `20/20` content-hashed partitions via `33/33` hardlinks and recomputed zero Parquet files; the mutation fixture rebuilds only the changed partition. Analytical gate passes. Promotion remains false because official totals reconcile only `83.8608%`, exact-label identities are not externally verified, and artifacts remain local-only. - The clean reusable semantic contract passes on the real accountability ledger at `126,760` entries after purging `10` fixture-derived money rows and their downstream issues/ledger entries: `13` typed partitions / `13` Zstandard files (`1,271,649` bytes), exact database/join/distinct-ID balance, `100%` public URL and evidence-lineage coverage, `126,734` source-record links, `126,757` resolved actor states, `3` explicit unresolved states, and zero private tokens. `362` legacy references use official source-default URLs without publishing workstation paths. The unchanged replay reuses `13/13` partitions via `13/13` hardlinks. This passes the real `S1` analytical gate, not million/public promotion: the lane remains parliamentary-dominated, local, and unpublished. - A separate deterministic accountability-ledger benchmark proves one-million transformation capacity without claiming real coverage or representative mix. It generates `1,000,000` source records/facts plus `1,000` issues at `11,222.877 rows/s` into a `1,454,428,160`-byte SQLite database (`quick_check=ok`, FK `0`, `44.719 MB` generation RSS). A source/year/entry expression index eliminates the temporary global sort. Export emits `20` Zstandard Parquet files / partitions (`9,908,795` bytes) in `60.063907s`; source URL, source-record, and lineage coverage are `100%`, actor states are `750,000` resolved + `250,000` explicit unresolved, unknown years and private findings are zero. Independent full-row/schema/hash/state validation passes; unchanged rebuild reuse is `20/20` through `20/20` hardlinks in `38.084129s`; process peak RSS is `367.266 MB`. Evidence explicitly sets `synthetic_capacity_only=true`, `real_coverage_claim=false`, and `representative_mix_claim=false`. Real `SCALE-034` remains open at `126,770` parliamentary-dominated facts. - A revision-preserving indicator lane and deterministic million-row benchmark close the synthetic indicator data-plane capacity gap without claiming real coverage. `indicator_observation_records` carries an additive FK to normalized series, and the public schema preserves source lineage, series/domain/geography, numeric/text/missing state, units, frequency, methodology, canonical dimension JSON plus SHA-256, and deterministic `sole` / `superseded` / `latest` revision state. Generation writes `1,000,000` observations, `5,000` series, and `10,000` source records into a `628,776,960`-byte SQLite database in `50.339267s` (`19,865.208 rows/s`, `quick_check=ok`, FK `0`). The corpus contains `950,000` logical groups and `50,000` truly revised groups. Its partition index avoids a temporary global sort. Export writes `16` source/geographic-scope/year partitions and `17` bounded Zstandard Parquet files (`41,200,775` bytes); independent validation checks every row, schema, checksum, partition hash/value, revision chain, lineage, URL, dimension hash, state, ordering, and private token. The unchanged rebuild reuses `16/16` partitions through `17/17` hardlinks with zero rebuilds; process peak RSS is `585.875 MB`. Evidence explicitly sets `synthetic_capacity_only=true`, `real_coverage_claim=false`, and `representative_mix_claim=false`. - The real indicator lane now passes the `S1` and `S2` row-scale analytical subgates in an isolated reproducible database. A bounded registry acquired four official Eurostat JSON-stat datasets—regional population by age/sex, regional GDP, regional population density, and regional poverty/exclusion—through verified TLS into content-addressed raw storage. Per-query byte and cube-cell ceilings fail closed before unbounded traversal; legacy queued payloads receive a conservative ceiling. The worker renews its lease during streamed download chunks and committed source-record batches, and aborts if ownership is lost. The durable queue recovered from the first parser failure and terminated `4/4` items with `0` unfinished/dead. Exact totals are `1,755,809` observations, `155,435` series, and `27,846,080` raw bytes. JSON-stat sparse values are traversed by bounded cube coordinates rather than sorting more than a million string keys. Source record IDs are content-versioned, while the raw-capture hash stays lineage-only so unrelated cube changes cannot create fake series revisions. - The first real backfill exposed a SQLite operational defect: commits inside a corpus-long read cursor retained an observed WAL near `16 GB`. Keyset pagination ordered by source/snapshot/record/PK plus `idx_source_records_indicator_backfill` bounds the replay WAL high-water mark near `496 MB` and keeps the query on the intended index. Dimensions now live once on `indicator_series` instead of being duplicated across `1,755,809` observations; observation/series raw payloads contain compact lineage pointers while full source payloads remain in `source_records`. After `VACUUM`, the database is `2,801,963,008` bytes. Replaying the backfill is idempotent at `155,435` series and `1,755,809` observations, with `quick_check=ok`, FK `0`, complete official URLs and methodology, zero inline observation dimensions, and exact acquisition/source-record/normalization reconciliation. - The real semantic export produces `26` source/geographic-scope/year partitions and `37` bounded Zstandard Parquet files totaling `253,373,860` bytes. The export completes in `252.681335s` at `272.875 MB` peak RSS. Independent validation reads all `1,755,809` rows and passes schema, file/partition hashes, min/max IDs, ordering, dimensions, revision state, lineage, official URLs, uniqueness, privacy, and a `1,536 MB` ceiling at `868.109 MB` peak RSS. This real run exposed two fixture-hidden bugs: generic file min/max metadata assumed identifier order matched sort order, and a zero revised-group total was treated as missing. Both are fixed with exact Arrow min/max computation and a zero-revision/non-monotonic-ID regression. Unchanged replay fingerprints the corpus and reuses `26/26` partitions through `37/37` hardlinks with zero rebuilds; full replay validation passes. Promotion remains false: the four datasets are valuable demographic/economic context, not representative policy-outcome coverage; there is no second-snapshot revision delta; and durable public-origin restore is unverified. - The actor/mandate backbone now passes the reusable local analytical contract at `79,105` mandates / `70,097` distinct people: `93` typed coarse-jurisdiction/source/year partitions and `93` Zstandard files (`8,193,685` bytes), with exact territory retained in-row. Database/join/distinct-ID balance, public URL coverage, and explicit lineage are `100%`; source-record coverage is `78,180/79,105`, `79,104` rows carry source-scoped identifiers, one row is explicitly `alias_only`, and private-token findings are zero. Export peak RSS is `152.734 MB`; independent full-row/hash/schema/list/identity-state validation peak RSS is `205.703 MB`. The unchanged next snapshot reuses `93/93` partitions via `93/93` verified hardlinks; the mutation fixture rebuilds one affected partition, and nested private identity values fail closed. This closes the local actor schema/partition/invalidation slice, not `S1` or promotion: the corpus is below `100,000`, source-scoped identifiers are not external identity proof, and the artifact is local/unpublished. - `RETIRED_NON_EVIDENCE`: the former public-money `S0` result used `10` fallback-derived facts and a v3 publisher that suppressed unclassified and potential-natural-person counterparties. It is retained only as failure history. No row, throughput result, redaction claim, or artifact from that run counts toward coverage, capacity, readiness, or publication. Current authority is the v5 identity-retaining contract over captured official data. - The new BDNS bulk acquisition lane discovered `28,669,785` live official concessions across `28,670` pages of `1,000`. It uses durable page jobs, bounded content-addressed raw downloads, page metadata/checksum lineage, batch upserts, crash recovery, pacing, and an adaptive upstream failure circuit. The `S1` acquisition cohort passes at `100/100` pages and exactly reconciles `100,000` distinct rows / `80,939,121` bytes with `0` pending/leased/dead pages. Immutable normalized fingerprints plus bulk-run/page-CAS/ordinal sightings now preserve future revision comparability without copying raw beneficiary payloads: the resumable backfill verifies every page checksum and reconciles `100,000` sightings, distinct IDs, versions, and raw-page links. Aggressive `16`- and then `4`-worker profiles initially triggered timeouts/resets; `40` interrupted leases were auditably reclaimed, then the final `45` pages completed without failure through one worker, two-page claims, and `2s` request-start pacing. Official public beneficiary names and source-published identifiers remain accountability evidence and must be retained exactly; only secrets, session state, workstation traces, and genuinely non-public data are restricted. This passes bounded acquisition/version-lineage `S1`; a second snapshot is still needed to prove real revision deltas, and it does not prove the full registry or promotion. - `RETIRED_NON_EVIDENCE`: the former 100,000-row live BDNS v3 semantic artifact suppressed `58,136` official-source unclassified counterparties. Its acquisition history remains useful, but that semantic root does not meet current publication policy and does not count. Regeneration must use v5, retain every source-published beneficiary name and identifier, and pass exact source-to-Parquet retention checks. - The prior fresh `S1` BDNS cohort replaced that missing/non-compliant root with official v5 evidence at `100,000` rows. It remains useful acquisition history but is superseded as current registry/readiness authority by the partitioned million-row `S3` cohort below. - Current BDNS authority is the `2026-08-12` partitioned `S3` cohort. Global deep-offset page 100 timed out even with one connection, so the durable queue uses official `fechaDesde`/`fechaHasta` daily windows, records partition lineage and source page numbers, and fails rows that escape their date window. The new `expand-daily` path revalidates official partition totals, preserves existing global page identities, and appends only missing source pages. It completed all `89/89` selected dates: acquisition finishes `1,419/1,419`, `0` retries/dead/unfinished, `1,080,788,680` raw bytes, and exactly `1,360,382` distinct source records, URLs, version sightings, and raw-page links. Every acquisition check passes. The v7 semantic artifact contains `1,360,382` rows in `14` Zstandard Parquet files (`42,955,289` bytes), carries the correct `s2_1m` capacity class, preserves the exact published award total `EUR 10,121,196,195.270000`, retains all `1,360,382` beneficiary names and all `163,270` source identifiers, retains `1,197,112` unclassified counterparties, and contains zero private tokens. The validator spills exact corpus-wide uniqueness to temporary SQLite instead of growing Python sets; canonical and unchanged-replay full-row validations pass. The unchanged replay reuses `1/1` partition through `14/14` hardlinks with zero rebuilds. The worker checks storage before every claim and reserves raw-object plus SQLite/WAL capacity. A temporary cleanup made one intermediate preflight green, but the latest captured acquisition preflight is `blocked_storage`: `5,685,862,400` bytes free versus `10,863,247,360` required, headroom `-5,177,384,960`; it claimed no additional work. Published HF release v2 `623b4a5a...` provides a verified durable analytical origin and enabled a separate clean-room drill from a new cache: all `20` selected files were downloaded and checksum-verified, then an isolated no-project environment with only copied validator code plus `pyarrow` full-validated all `1,360,382` rows at `350.484 MB` peak RSS. This closes selected-window acquisition, real million-row capacity, durable analytical origin, and analytical clean restore. Full history, second-snapshot revisions, representative completeness, rollback rehearsal, and acquisition storage headroom remain open. - The ordinary Hugging Face tabular snapshot for `2026-08-12` is published and remotely verified at `127` tables / `134` Parquet files; its red vote and initiative quality gates remain explicit. Published content-addressed analytical release v2 `5872efaf...` is remotely verified for all six registered corpora: `5,403,506` rows, `8,595` canonical data files and `498,631,274` bytes, preserving every official public identity. Pointer, manifest and stable artifact contract `bb99c119...c2c` retain exact data parity. The current live check has zero errors and two metadata-freshness warnings because local registry/readiness now include the later document-eligibility audit and no republish is authorized. Two fresh empty caches restored all six corpora: BDNS downloaded `20` selected files / `43,042,223` bytes, while the other five downloaded `8,601` / `461,088,883`; every byte passed SHA-256 verification. Isolated copied validators full-read every row and data file with bounded memory and zero private-token findings. All six `durable_public_origin` and `clean_room_restore` flags are true. Authorized pointer activation, database rebuild from immutable inputs, and the raw-object bucket contract remain open. - Explicit immutable-release recovery no longer depends on `scale/latest.json`: `--snapshot-path scale/snapshots//` validates the target against the fetched manifest before downloading. A full fresh-cache rollback-candidate drill against prior published release `623b4a5a...` downloaded and checksum-verified `8,619/8,619` files / `504,093,303` bytes in `1,199.614s`, then full-validated all six corpora / `5,403,506` real rows and official-public identity retention with `1,007.453 MB` peak RSS. The disposable restore cache was deleted after retaining evidence. A GET-only rollback plan then re-fetched live current `5872efaf...` and target `623b4a5a...`, verified both content-addressed manifests, bound restore/full-validation/live reports by SHA-256, proved a zero row/file/byte contract delta, and emitted compare-and-swap plus supersession requirements that preserve both immutable releases. It records `authorized=false` and `mutation_performed=false`. Authorized pointer activation, formal RPO, and activation RTO remain open. - The restored actor lane now rebuilds through bounded batches into an atomic SQLite artifact. Two independent runs both contain `88,031` rows, `88,031` distinct mandate IDs, and exact logical-row hash `eb7fdb8e...`; their `71,168,000`-byte SQLite outputs are byte-identical at SHA-256 `61cfdf8e...`, with integrity and memory gates passing. This proves deterministic analytical reconstruction while leaving normalized production-schema and relationship reconstruction open. - Raw-object replication and restore are bounded-parallel instead of single-threaded. Two 16-worker local CAS replays each reconcile the older `6,792`-object linked-text subset / `133,219,457` bytes, deduplicate every object, and emit the same deterministic manifest; a streaming `full_manifest` drill restores and checksum-verifies all `6,792`. These are warm local contract/throughput results, not evidence of remote S3 durability or coverage of the current `111,408`-document inventory. The current native-text CAS contains `108,139` checksum-verified objects / `533,125,716` compressed bytes; remote replication and clean restore of that full set remain open. - Document provenance now reconciles `111,408/111,408` eligible real files and covers `111,357` with checksum plus public URL (`99.9542%`). Existing exact evidence classes remain separate: embedded official URLs, captured hrefs plus byte-identical GETs, archived SQLite metadata, official publication identifiers, official path projections, explicit sidecars, real manifests, and published JEC/Andalusian Parliament indexes. All `89,001` BOE documents and `1,034` summaries are linked. The changed historical Senate response remains rejected. The residual remains exactly `51` files / `8,315,544` bytes with zero checksum conflicts and no invented URLs. - Immutable archive recovery is now a separate disk-backed lane. It derives only two exact source URLs from current connector/path evidence, covers all `26` Europarliament/Senate gap instances as `3` SHA-256/size targets, and persists source queries plus replay outcomes in SQLite. The first bounded run found `5` unique older Europarliament archive digests, streamed every replay, and rejected all `5` because neither bytes nor SHA-256 matched the current preserved targets. Seven endpoint variants ended in timeout or HTTP `503` and remain explicitly retryable; successful index queries and rejected digests are not repeated by default. The manifest remains empty, the v12 provenance audit loads that empty exact-match contract, and no gap is closed from URL resemblance alone. - The residual is operational rather than narrative: a deterministic streaming export validates the audit/file-manifest hash contract and emits `51` remediation tasks over `8,315,544` source bytes. Every task lacks both checksum lineage and public URL; lanes and priorities are explicit, and guessed URLs remain forbidden. The remaining groups are one nonmatching historical Senate session XML / `754,640` bytes, `25` Europarliament captures / `4,398,347` bytes, and `25` primary-party program captures / `3,162,557` bytes. - The scale-readiness release gate now consumes seven registered real-corpus contracts plus inventory, provenance audit v12, gap queue, checkpointed Wayback report, `43`-row portable checkpoint, exact-match manifest, native extraction, independent native validation, page OCR, independent OCR validation, the human assignment contract, review validation, and the reviewed CSV as typed evidence. Candidate readiness independently checks normalized BOE lineage, exact public-name/source-text retention, correction-backed source records, null outcome semantics, and unchanged replay. The gate fails on schema drift, stale counts, archive target/path/query/candidate/queue/audit imbalance, broken raw/CAS hashes, any native truncation, incomplete page balance, immutable-assignment drift, incomplete reviewer fields, identity-retention policy drift, or invented human fields. Review preparation refuses to overwrite completed human fields. The refreshed full-row readiness run remains `real_foundation_ready_scale_incomplete`: seven real corpora validated, three million-row lanes, zero promoted. - Final local verification passes `1,346` Docker-backed tests with `11` contextual skips, audits all `92` staging databases with zero integrity findings, scans `124,418` real-data artifact files with zero synthetic/mock findings, scans `16,447` public artifacts with zero secret/session/workstation findings, and passes repository-root hygiene. Official public-domain identity fields remain retained by policy and by the machine-checked artifacts. - Public-safe storage health now consumes only the real BDNS and PLACSP preflight contracts, rejects unknown schemas or contradictory arithmetic/state, hashes each evidence source, and omits roots, credentials, sessions, and workstation paths. The local publishable artifact is `blocked_storage`: BDNS needs `5,177,384,960` more bytes and PLACSP `66,776,911,872`. It is wired into the next authorized HF snapshot package but has not been published remotely in this run. - A separate `S2` cohort has `1,000` bounded pages durably enqueued. Its immutable checkpoint contains `146/1,000` pages, `146,000` exact distinct rows, `114,558,806` bytes, `146,000` checksum-linked version sightings, `0` dead pages, and `854` pending. The run exposed that a per-claim circuit could stop on one isolated failure; the worker now has a tested rolling `20`-outcome minimum window. The corrected profile then stopped honestly at `16` upstream failures in `24` attempts. This source is `no_new_lever` until fresh upstream/session evidence; the queue remains recoverable. This is progress evidence, not a passed million gate; official-source identities remain retained evidence. - The acquired `146,000` rows continue through the v3 semantic pipeline without upstream access: one partition, two Zstandard Parquet files (`3,271,529` bytes), exact award total `EUR 3,473,893,337.060000`, complete source/lineage/amount/EUR/nonnegative/privacy coverage, `43,583` published legal entities, `102,417` withheld unclassified counterparties, and zero private tokens. Independent full-row/schema/hash/decimal/semantic/privacy validation peaks at `450.609 MB`; unchanged rebuild reuse is `1/1` through `2/2` hardlinks. This validates the partial S2 artifact, not one-million capacity or public-origin promotion. - `RETIRED_NON_EVIDENCE`: the former one-million generated public-money benchmark is historical engineering context only. Generated rows and its suppression policy are forbidden as coverage, capacity, readiness, or publication evidence. Only real official corpora now count. - A separate deterministic actor-mandate benchmark now proves one-million transformation capacity without claiming real identities or coverage. It generates `1,000,000` source records, people, and mandates plus `800,000` source identifiers and `100,000` aliases into a `1,250,463,744`-byte SQLite database in `59.948452s` (`quick_check=ok`, FK `0`, `40.016 MB` generation RSS). The first scale attempt exposed a pre-write global sort/materialization long pole. Covering actor indexes plus correlated per-person rollups reduce the measured million-row fingerprint pass from more than seven minutes without completion to `35.567450s`. The final export emits `20` Zstandard Parquet files (`19,728,137` bytes) in `60.916538s`; independent full-row/schema/hash/list/identity/privacy validation passes; unchanged rebuild reuse is `20/20` through `20/20` hardlinks in `35.338849s`; process peak RSS is `507.531 MB`. Evidence explicitly sets `synthetic_capacity_only=true`, `real_coverage_claim=false`, and `external_identity_verified=false`. Real `SCALE-032` remains open at `79,105` mandates. Evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/document-pipeline-scale-run.json`. Format evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/real-document-format-inventory.json`. OCR routing evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/pdf-ocr-routing-benchmark.json`. Complete native extraction evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/real-document-text-extraction.json` and `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/real-document-text-validation.json`. Complete page OCR evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/real-document-page-ocr.json` and `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/real-document-page-ocr-validation.json`. Human OCR review packet evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/real-document-page-ocr-human-review-queue.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/real-document-page-ocr-human-review-validation.json`, and `docs/etl/sprints/SCALE-FOUNDATION-20260810/exports/real-document-page-ocr-human-review.csv`. Vote source URL evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/member-vote-source-url-lineage-20260812.json`. Senate repair evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/senado-local-cache-repair-audit.json`. SQLite-direct shard evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/senado-repaired-db-shard-manifest.json` and `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/senado-repaired-db-shard-validation.json`. Semantic partition evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/member-vote-semantic-partition-manifest.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/member-vote-semantic-partition-validation.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/member-vote-semantic-partition-incremental-manifest.json`, and `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/member-vote-semantic-partition-incremental-validation.json`. Accountability-ledger partition evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/accountability-ledger-semantic-partition-manifest.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/accountability-ledger-semantic-partition-validation.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/accountability-ledger-semantic-partition-incremental-manifest.json`, and `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/accountability-ledger-semantic-partition-incremental-validation.json`. Accountability-ledger million-capacity evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/accountability-ledger-1m-capacity-benchmark.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/accountability-ledger-1m-capacity-manifest.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/accountability-ledger-1m-capacity-validation.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/accountability-ledger-1m-capacity-incremental-manifest.json`, and `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/accountability-ledger-1m-capacity-incremental-validation.json`. Indicator million-capacity evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/indicator-observation-1m-capacity-benchmark.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/indicator-observation-1m-capacity-manifest.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/indicator-observation-1m-capacity-validation.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/indicator-observation-1m-capacity-incremental-manifest.json`, and `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/indicator-observation-1m-capacity-incremental-validation.json`. Real Eurostat indicator evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/eurostat-indicator-real-s2-acquisition.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/eurostat-indicator-real-s2-semantic-manifest.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/eurostat-indicator-real-s2-semantic-validation.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/eurostat-indicator-real-s2-incremental-manifest.json`, and `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/eurostat-indicator-real-s2-incremental-validation.json`. Actor-mandate partition evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/actor-mandate-semantic-partition-manifest.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/actor-mandate-semantic-partition-validation.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/actor-mandate-semantic-partition-incremental-manifest.json`, and `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/actor-mandate-semantic-partition-incremental-validation.json`. Public-money partition evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/public-money-semantic-partition-manifest.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/public-money-semantic-partition-validation.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/public-money-semantic-partition-incremental-manifest.json`, and `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/public-money-semantic-partition-incremental-validation.json`. Current BDNS bulk evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/bdns-concessions-partitioned-real-s3-enqueue-20260812.json` and `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/bdns-concessions-partitioned-real-s3-run-20260812.json`. Current BDNS v7 semantic evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/bdns-public-money-real-s3-v7-manifest-20260812.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/bdns-public-money-real-s3-v7-validation-20260812.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/bdns-public-money-real-s3-v7-incremental-manifest-20260812.json`, and `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/bdns-public-money-real-s3-v7-incremental-validation-20260812.json`. HF origin evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/hf-public-origin-quality-20260812.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/hf-scale-origin-bundle-dry-run-20260812.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/hf-scale-origin-publish-20260812.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/hf-scale-origin-verify-20260812.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/hf-scale-origin-clean-restore-bdns-20260812.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/hf-scale-origin-clean-restore-bdns-validation-20260812.json`, and `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/hf-scale-origin-clean-room-drill-20260812.json`. BDNS million-cohort evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/bdns-concessions-s2-enqueue.json` and `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/bdns-concessions-s2-partial-run.json`. BDNS partial-S2 semantic evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/bdns-public-money-semantic-s2-partial-manifest.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/bdns-public-money-semantic-s2-partial-validation.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/bdns-public-money-semantic-s2-partial-incremental-manifest.json`, and `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/bdns-public-money-semantic-s2-partial-incremental-validation.json`. Public-money million-capacity evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/public-money-1m-capacity-benchmark.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/public-money-1m-capacity-manifest.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/public-money-1m-capacity-validation.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/public-money-1m-capacity-incremental-manifest.json`, and `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/public-money-1m-capacity-incremental-validation.json`. Actor-mandate million-capacity evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/actor-mandate-1m-capacity-benchmark.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/actor-mandate-1m-capacity-manifest.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/actor-mandate-1m-capacity-validation.json`, `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/actor-mandate-1m-capacity-incremental-manifest.json`, and `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/actor-mandate-1m-capacity-incremental-validation.json`. Document-manifest million-capacity evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/document-manifest-1m-capacity-benchmark.json`. Document-fetch local HTTP fault-matrix evidence: `docs/etl/sprints/SCALE-FOUNDATION-20260810/evidence/document-fetch-local-http-fault-matrix.json`. ## Where we are going - Replace the local origin with a versioned remote S3-compatible origin and repeat full-manifest restore from a disposable cache. - Turn the numeric `111,408`-document cohort into a representative quality cohort across digital PDFs, scanned PDFs, office formats, large/malformed files, languages, and blocked sources; then grow it to one million. - Finish human OCR review and publish quality, cost, and throughput by stratum before claiming document-scale promotion. - Keep old Senate detail-route retries paused until a fresh session/cookie, stable official endpoint, or archive source provides a new lever. ## What is next 1. Review the eight captured historical BOE correction documents and convert their `776` exact paragraphs into reconstructable operations or explicit context; keep all `31,241` affected base candidates out of effective state until each cohort reconciles exactly. 2. Activate and revert the prepared HF rollback plan only under explicit authority, record supersession/RTO, approve formal RPO, and separately finish the raw-object S3-compatible bucket/retention contract. 3. Complete all `40` real OCR review decisions, reviewer fields, and accuracy/usability summaries without machine-invented labels. 4. Add representative scanned PDF, office, multilingual, malformed, large-file, and multi-institution cohorts to the passed `111,408`-document manifest; reuse the proved native/OCR workers and quarantine, then grow through bounded official-host/date cohorts to one million. 5. Reconcile `discovered = succeeded + dead + deferred` per source, legislature and route in every release. 6. Recover the missing authoritative member rows behind `1,115` classified discrepancies, resolve Senate identities externally, then publish the already validated SQLite-direct shard set to verified origin/CDN. 7. Capture a second real BDNS snapshot/revision delta, extend history, then close beneficiary resolution. Use the proven actor path to grow the real actor corpus past `100,000` and then one million with external identity adjudication; add representative PLACSP; expand the passed `1,755,809`-row Eurostat cohort into a representative Tier-1 outcome/confounder mix; expand ledger from `S1` to representative `S2`. ## Visible progress gate `PASS`: real evidence advanced to complete native processing of `111,408` document instances / `109,558` contents and complete additive OCR of `856` routed pages, both independently validated with zero native truncation and all official public identity retained; human review remains `0/40`. BOE now preserves `42,056` exact base candidate occurrences across seven proclamations. The reviewed subset applies three correction documents through `17` reconstructable operations to produce `10,820` effective rows; eight correction documents and their `31,241` candidates remain explicitly pending and excluded. Their `776` exact paragraphs now have deterministic blank dual-review tasks backed by one deduplicated full public-candidate context corpus and enforced through central readiness. The two-file semantic artifact preserves all effective exact names/text/evidence, keeps results unknown, independently validates, and replays `2/2`. Member votes pass at `1,809,222`; accountability ledger at `126,760`; PLACSP at `263,302`; actor mandates at `88,031`; BDNS at `1,360,382`; Eurostat at `1,755,809`. Seven real corpora validate; three exceed one million; zero are promoted. No synthetic or mock record counts. Residual correction decisions, source defect, human review, upstream, identity resolution, publication, representation/history, revision-delta, real coverage, and reconciliation gates remain explicit.