Disc coverage
How much of the game's own bytes the project can account for, measured against the disc rather than against the project's own notes. This is where the byte-denominated figures on the landing page come from - and the one instrument that can say how much of the game is left.
At a glance
- Tool
scripts/ci/disc-coverage.py; attribution sweepscripts/ghidra-analysis/attribute-dump-extents.py- Two halves
- Code is byte-exact (a byte is inside a dumped function or it is not); data is format recognition, an upper bound
- Inputs
- Ghidra dump corpus (
ghidra/scripts/funcs/) +extracted/PROT/categorize.json- both gitignored, disc-derived - Gate
--checkratchets againstscripts/ci/disc-coverage-baseline.json; skips and passes without disc data- By-product
- The largest un-dumped code runs, listed as a dump worklist derived from the bytes
- Sibling
- Port catalog - the same questions denominated in what the project cites
What this solves
The port catalog tracks dumped, documented and ported over the set of addresses this project cites - the right instrument for steering work, but blind to a subsystem nobody has cited. disc-coverage.py takes the denominator from the disc instead. Reach for it when you want to know what code nobody has looked at yet, or when you want a progress figure that cannot flatter itself.
The distinction runs in both directions: closing a byte-denominated gap widens the citation graph, because a newly dumped and documented function becomes a row on the port worklist. A rising port worklist after a dump pass is the measurement getting wider, not the work going backwards.
Quick start
python3 scripts/ci/disc-coverage.py # report -> target/disc-coverage/
python3 scripts/ci/disc-coverage.py --md # markdown to stdout as well
python3 scripts/ci/disc-coverage.py --check # ratchet against the baseline
python3 scripts/ci/disc-coverage.py --update-baseline
The data half reads extracted/PROT/categorize.json. That file is a cache written by asset categorize and never regenerated automatically, so a tree whose detectors have moved on keeps reporting a stale classification - and the gate passes either way. Regenerate before trusting a data figure or taking a baseline:
./target/release/asset categorize extracted/PROTThe two halves measure different things
| Kind | What a percentage means | |
|---|---|---|
| Code | byte-exact | a byte is inside a dumped function, or it is not |
| Data | format recognition | an entry's format class is known; its bytes are not individually accounted for |
The data figure is an upper bound: knowing an entry is a scene_vab_stream is not the same as consuming every byte inside it, and no parser in the tree reports consumed-versus-unconsumed extents. The data denominator also counts some disc bytes more than once - entry weights use max(indexed_size_sectors, footprint_sectors), and only the footprint is an entry's real extent (see PROT TOC) - so the totals run to roughly 2.5× the archive's real size. Quote the shares, not the totals; when a single entry's weight makes a residue look significant, check its footprint first.
A statistical class is not a verdict
asset categorize ends in a statistical fallback that buckets unclaimed entries by zero fraction and entropy (mostly_zeros, unknown_low_entropy, ...). When reading the per-class table:
- A statistical class name is the absence of a finding, never a finding -
mostly_zerosmeans "sparse and unclaimed", not "empty". - The placeholder column is only as trustworthy as the detectors that ran before the fallback; a format with no detector is invisible to it.
- Entries of an identical, exact size are the cheapest lead: a size repeating across a hundred entries is a fixed-layout format whatever its statistics look like.
- A statistical class can swallow content by dilution - a whole-buffer printable-ASCII ratio misses an image that is mostly bss with its literals in the first sector. Structure at a known offset beats a whole-buffer ratio.
History: the field maps that hid inside "placeholder"
Every scene's field map - collision grid, floor heights, object placements, door triggers - is a sparse 0x12000-byte blob mostly zero on a small map. Before the field_map detector existed, those entries scattered across the statistical classes by nothing but how crowded each scene happened to be, and the largest single block of them sat inside the "explained placeholder" column. That is the inference being wrong at scale, and the reason the bullets above are phrased as they are.
How code coverage is computed
Every Ghidra dump header carries an entry address and a byte length (== FUN_800402f4 ... size=7904 bytes, 1976 instructions ==), so the dumped functions are real intervals over an image's address space. The script merges them, subtracts them from the image's extent, and splits the remainder into code and data - a PS-X EXE's text segment carries read-only data in the same span, and counting string tables as "un-decompiled code" would understate coverage badly. Each gap is classified statistically (share of plausible MIPS opcodes, density of 0x80xxxxxx pointer words), checked against a known-code control region.
A code gap is not automatically un-analysed code
The report classifies each code gap by shape, and only one of them is work:
| Shape | What it is | In the code denominator |
|---|---|---|
code | genuinely un-dumped instructions - the work | yes |
data | the opcode statistic rejects it: rodata in the text segment | no |
padding | every word is nop: inter-function alignment | no |
mostly_padding | at least half the words are zero | no |
no_exit | 1024 bytes or more with no jr ra in them | if the statistic passes it |
return_tail | a jr ra (+ nop) the preceding routine's analysed body stops short of | if the statistic passes it |
bios_thunk_slot | the delay slot of a jr $t2 PSX BIOS-call thunk | yes (tiny-gap fiat) |
psyq_lib_stamp | an 8-byte librarian version stamp between link modules | yes (tiny-gap fiat) |
constant_table | one repeated non-nop constant: a data table resident in the text segment | yes (tiny-gap fiat) |
The end-of-body and linker shapes persist however much is dumped, which is why the figure asymptotes short of 100%. Most shapes are reported, not subtracted - the denominator keeps them, so the ratcheted figure stays comparable across classifier changes. The padding shapes are the exception, excluded by a structural rule rather than by the opcode statistic: a word of zeros decodes to nop, so a majority-zero run passes the code test outright and would otherwise sit at the top of the worklist while the census called it padding.
The same encoding defeats the attribution sweep, one level up. That sweep names the image an ambiguous extent belongs to by asking which image's own bytes reproduce the dump's opening instruction window at that address - and a window of nops reproduces inside any image's zero fill, at any base. A window's usable length is therefore its non-nop instruction count. Two shapes of false credit came out of the corpus when that guard was added: an all-nop window read as byte-identical across two images that agreed about nothing but being empty, and a 24-instruction window carrying one lone non-nop word in a run of zeros read as a unique attribution, crediting one image with a 20 KB extent that is over nine-tenths zero fill. Removing the credit moved most images' floors up - extents the zero windows had made ambiguous for several images stopped being ambiguous at all - and only the images the zero fill had been credited to down.
What the executable's gap turned out to be
Working the SCUS_942.54 gap list until it stopped yielding produced on the order of a hundred function entries. A five-form reference sweep (address-reference scan) over them splits three ways:
| Share | Why it had no dump | |
|---|---|---|
| no reference of any form | ~3/4 | Ghidra creates functions from the call graph; a routine nothing references gets no function record, no dump, no citation - invisible to any citation-denominated instrument. |
referenced from SCUS_942.54 | ~1/5 | mostly reached through the entry-stub / init path, or from a body whose own analysis was incomplete. |
| referenced only from an overlay | a few | live code whose caller lives in an image the SCUS-only analysis never sees - "zero static callers is not dead", showing up as a coverage gap. |
The first row is the point: the bytes find the class the citation graph structurally cannot. Caution: "no reference exists" is a statement about static references, and much of that band is linked-in library material the game does not call.
Dump exclusions: a census, not a count
The dump corpus spells every header field more than one way, so the report names each reject class and separates two populations: answers (pointer_stub, nofunc_record, data_window, not_a_dump - the corpus recording a result, not defective and not work) and defects (zero_insns, gapped_stream, no_extent, empty_dump - a dump that cannot evidence its own extent). A pointer stub is the corpus doing the right thing with a mid-function address; zero_insns is Ghidra's "bad instruction data" - a data window asked for as code, excluded rather than credited.
Reading the header looked like the trivial part and was the part that went wrong: a dozen dump scripts each spell the header differently, and each instrument once grew its own regex, silently rejecting a different set of real dumps and reporting them as a corpus deficiency. Sharing one parser (scripts/ghidra-analysis/dump_header.py) across every instrument is the fix; the general lesson is on dump-corpus-integrity.
Why overlay rows can read "not meaningful"
SCUS_942.54 is the only image with an unambiguous answer: one load image, one fixed base. Several overlays load at the same base (0x801CE818), so a dump whose entry lands in that band cannot be attributed to one image by address alone. Rather than publish a number that quietly double-counts, each overlay row carries the share of its covered bytes that could not be placed; above 50% the figure is replaced by not meaningful and the row leaves the ratchet baseline. That is also why the landing page's decompilation bar quotes the executable alone.
The ambiguity is resolved where the bytes can resolve it: attribute-dump-extents.py disassembles each extracted image at its static-overlays.toml base and asks which images actually hold a dump's bytes at the VA it prints. Its per-extent verdict is committed as scripts/ghidra-analysis/dump-extent-attribution.csv:
| Verdict | Meaning | What the gate does |
|---|---|---|
unique | one image holds those bytes there | credit only that image |
identical | several hold byte-identical code there | credit each of them |
misbased | the bytes live at another VA entirely | credit nobody |
gapped / data | not a coherent function body at that VA | credit nobody |
short / unresolved / no_disassembly | the window cannot sign it | residue: stays ambiguous |
The key is the extent (entry, bytes), not the dump filename, so the file does not rot when a dump lands or is renamed. Attribution is optional: without the CSV every overlay extent stays ambiguous - the honest upper bound, just a looser one.
What the residue is, and which count to quote
| Residue shape | What would close it |
|---|---|
| a few-instruction window no image reproduces at that VA | nothing cheap - too short to search for elsewhere without inviting a coincidental hit |
| bytes in no extracted image at any VA | an extraction, not a dump - live-RAM captures of overlays never statically extracted; the static overlay pipeline's job |
| two dumps at one extent resolving to different images | already answered: several routines share the range - a finding, not a gap |
The sweep asks two questions with opposite sensitivities to window length - does this image reproduce the window at this VA (fixed-offset test, three-instruction floor) and do these bytes appear anywhere (a search over millions of positions, eight-instruction floor); --validate-short-floor sets the relaxation from its own control. The code table carries three counts, and the unit matters: dumps is per dump file; VA-ambiguous is the share of an image's covered bytes the attribution could not place, which is the unit the coverage figure itself is stated in; by extent is the same question counted in whole extents. Reading the extent count as the discount on a byte figure withheld an upper bound from most of the slot-B band over fractions of a percent of its bytes - the corpus carries a long tail of 4-to-36-byte Ghidra fragments, and an extent count weighs each of those like a 6 KB module. Neither is per dump file: one extent can back dozens, and weighting by how often the same bytes were dumped measures the corpus, not the image.
History: the starting point mistaken for the limit
This page once concluded that the inner of two nested overlay spans could never be measured at all. That reading is falsified: address ambiguity is total for the inner span, but byte attribution then places most of its extents - what is structural is that it cannot be resolved by address, a statement about one method, not about the image. A second earlier claim, that the residue is mostly repaired by re-dumping, is also unsupported. Both are logged in do-not-re-walk; companions dump-corpus-integrity and phantom-print-index.
Gate behaviour
- Both inputs are gitignored, so a clone without disc data exits 0 and reports SKIPPED - the same skip-and-pass convention as the
LEGAIA_DISC_BINtests. The ratchet only has teeth on a machine with the disc. --checkcompares against the committed baseline; coverage may only go up, within half a percentage point. A baseline moving down is a claim that needs a reason in the commit message (--update-baseline).- The report lists the largest un-dumped code runs in
SCUS_942.54by size - the one dump worklist the citation graph structurally cannot produce. - The landing-page figures are refreshed by
scripts/ci/update-progress-metrics.pyand committed, because the site is built where the disc is not.