At a glance

Tool
asset account <entry>; whole-TOC sweep scripts/asset-investigation/byte-account-sweep.py
Module
legaia_asset::byte_account
Question
Not "what is this entry" but "which of its bytes does anything here consume, and what is the rest"
Two figures
structural (a walker followed the container's own layout) and accounted (that, plus what a magic sweep found in the residue)
Output
Per-owner claims, per-shape residue, and the largest unclaimed runs - a worklist
Sibling
Disc coverage - format recognition, weighted by entry size, and an upper bound by construction

Running it

asset account extracted/PROT/0867_battle_data.BIN
asset account 893 --prot-dir extracted/PROT      # bare extraction index
asset account 898 --funcs ghidra/scripts/funcs   # an overlay code image
asset account 0086 --depth 2 --json > kingdom.json

The method is four steps. Classify the buffer with asset categorize and pick a walker; run the walker, which emits claims - half-open [start, end) ranges each tagged with an owner; merge the claims and take the complement, so every maximal uncovered run is one residue; and where a claim covers a compressed span, decode it and account the payload in a nested pass. Two walker selections are keyed on the extraction index rather than on the class, because no detector fires on their heads: the monster archive, and any entry with a row in the static overlay map.

What the number means

accounted is the share of an entry's bytes some parser consumes. It is not the share whose meaning is understood: a fixed-region file accounts to 100% the moment its region constants are written down, while the semantics of those regions can be only partly pinned. Read it as an upper bound on understanding and a lower bound on structure.

Claims are tiered so the figure cannot be inflated by guessing. An owner of scan means a magic sweep found the sub-asset in the residue - evidence that it is there, not that anything walked to it. That is why the report prints structural beside accounted, and why a large gap between them is a warning rather than a result.

Owner vocabulary

An owner says what kind of thing consumes the bytes, not which module claimed them; each claim's free-form detail names the instance.

OwnerBytes it names
header / toca container's fixed head words; the offset or descriptor table its reader walks
lzsa compressed span - the decoded side is accounted in a nested pass
tim / tmd / vab / seq / anma sub-asset of that format, header through end
record / grida fixed-stride data record; a dense per-tile grid
script / codea bytecode body; MIPS instructions inside a dumped function's extent
texture / clut / stringa raw texture page with no TIM header; a palette region; a resolved string
paddeclared slack inside a fixed-stride slot the container's own size math covers
scanfound by a magic sweep over the residue, not by a structural walk

Residue shapes

Each uncovered run gets exactly one shape, and the tests run in a fixed order that is load-bearing.

ShapeTest
zero_padevery byte is 0x00
alignmentshorter than 16 bytes and not all zero
repeated_fillone short pattern (period 1, 2, 4, 8 or 16) repeated - dev fill
ascii_textat least 80% printable ASCII or NUL - a string pool
pointer_denseat least 20% of words land in 0x80000000..0x80200000
bgr555nearly every halfword below 0x8000, wide high-entropy spread - PSX 15-bit colour
plausible_mipsplausible opcodes over at least four of them, almost no pointers, no SPECIAL landslide
low_entropy / high_entropy / mixedbelow 4.0 bits/byte; at or above 7.2; none of the above

Three ways a naive classifier calls data "code"

  • A pointer table decodes as code. A word like 0x801C1234 has primary opcode 0x20 - a plausible lb - so a table of overlay pointers scores perfectly on opcode plausibility. pointer_dense is tested first.
  • A small-value table decodes as code. A sparse 16-bit table is SPECIAL in every word, which is also plausible. So plausible_mips additionally demands a spread across four primary opcodes and keeps SPECIAL under 60%: real code mixes loads, stores, immediates and branches; a table does not.
  • A colour page decodes as code too. A BGR555 halfword pair forms a plausible word and colour spreads across enough opcodes to clear the distinct-opcode floor, so bgr555 is tested first - and it discriminates the other way, because real MIPS puts a halfword at or above 0x8000 in every load, store and lui pair.

Overlay code images

For an entry that is a runtime overlay, the "parser" is the Ghidra dump corpus: a dump header states an extent, and that extent maps to file offsets through the overlay's load base. The header parse is re-implemented in Rust, so the instrument needs a checkout and a dump directory and never invokes Ghidra.

The hard part is not the header. Every slot-A overlay loads at 0x801CE818, so a printed address cannot say which image a dump belongs to. This instrument resolves it from the bytes: it re-encodes the dump's first printed instructions and compares them against the image at the mapped offset. A printed instruction that re-encodes to the image's word is confirmed and credited; one that re-encodes to a different word is refuted and dropped; a dump whose head is outside the small encodable grammar is unverifiable and is credited only when its filename label names this entry. The grammar is deliberately small, so the instrument's silence never reads as a refutation.

The slot-B images

An image walks this way when its map row is linked at the slot-B base - not when its index falls in a range, and not on a class. The three regions the walk recovers are each resolved by comparing a word against that base, so the base is what makes the walk apply, and it runs with or without a dump directory because those regions come out of the image itself. The walker claims the head jump table as toc and each spawn record as record.

Seventy mapped images sit at that base. Sixty-four are the cast / summon band 0903..0966 - the per-spell summon stagers and capture-class cast modules the battle overlay's three entry tables reach. The other six are the two render occupants, the battle tutorial and its two stage-module siblings, and the staged texture loader. Selecting on the index band had left all six measured as if they had no head table.

A record's extent is named by an instruction, not by a statistic: the module materialises the address with a lui/addiu pair and hands it to FUN_80021B04 / FUN_80050ED4 in $a2, and the next record's pointer is the current one's end. The image's highest record has nothing above it to bound it, so its own move-VM program's terminator does - a HALT, an armed idle loop, or, only where neither turns up, the last maximal WAIT the walk stepped over. On retail every image that has a highest record bounds it; where a walk does not terminate the record stays residue and the walker's note names its offset. Full layout: slot-B module layout.

That makes the set the cleanest class on the disc, and the numbers say why in a way no other class can match. All 70 images select this walker over 653312 bytes, and every claim is structural - not one byte of the accounted share comes from a magic guess. The residue is 11216 bytes, of which 2618 are zero padding. Per entry the accounted share runs 69.2% to 100.0% with a median of 99.1%; the floor is the battle tutorial, whose own content above its code is a string pool no consumer has been traced to yet.

Reading a report

  • A large high_entropy run nothing claims is a compressed or sample-data region with no walker. Top of the worklist.
  • A large plausible_mips run in a non-overlay entry is either an un-dumped routine or a walker that stopped early.
  • A low_entropy or mixed run next to a claimed table is usually that table's body: the walker found the header and stopped.
  • zero_pad and repeated_fill are not work - sector padding and dev fill are the disc's own slack.
  • A high structural with a small gap to accounted is the healthy case. The reverse means most of what was found was found by guessing at magics, and the container's layout is still unwalked.

Ranked residue across the whole TOC

One entry at a time is the right shape for a hunt and the wrong shape for a worklist: the entries whose residue is worth walking are not the ones anybody thought to run the tool on. The sweep runs every PROT entry and rolls the result up by entry class and by residue shape, so what lands at the top is a property of the disc rather than of what somebody sampled. It writes a CSV of numbers only - the per-run head hex is disc bytes and is dropped - under target/, and nothing about that file is committed.

ClassWhat its unclaimed bytes areVerdict
the fixed-stride streaming slotsnothing - each slot's fill is claimed out to the stride the loader transfersclosed; the monster archive alone used to be four fifths of the disc's residue
the multi-bank VABnothing - the bank index, each bank's two chunks and each bank's sector slack are all claimed from lengths the container statesclosed; it used to be the largest non-slack residue on the disc
overlay_data_blobalmost all zero_pad; the rest is one entry whose whole extent reads plausible_mips although its head is a length-prefixed Japanese label tablea Japanese-build menu image the USA disc never loads
lzs_containerone member that is a code image rather than a container; the per-entry tails past the last descriptor's stream were the packer's buffer and are claimedone mis-classed entry
overlay_ptr_table, mips_overlaylow_entropy with a plausible_mips minoritydump worklist - and it agrees with disc coverage's gap list
scene streams, battle datanothing - the chunk walk reaches the terminator and what is left of the entry's last sector is claimed as slack, the same thing the multi-bank VAB's per-bank slack isclosed; this used to be the largest single figure in the ranking, under the verdict "walker tails"
scene asset tables, packsnothing but zero_pad - the one run per bundle that used to rank here is an earlier entry's bytes at the same file offsets, above the last descriptor's streamclosed; "walker tails" was the wrong verdict - the walker had reached the end of the content
filler, zero entries, trigger sidecarsascii_text and zero_padthe disc's own slack
field mapsnothing - every entry accounts fullyclosed

Two ways the headline number lies

A low percentage that is finished. Every walk-on trigger sidecar accounts in the single digits and every unclaimed byte in all of them is zero_pad: the format is one sector whose records occupy a small head and whose rest is padding it declares. Ranking a worklist by accounted share alone would put all of them near the top. Rank by non-slack residue instead.

A high percentage that walked nothing. The multi-bank VAB used to account for almost all of its bytes and structurally for none of them - every claim came from the magic sweep, which is evidence that sample bodies are there and nothing more. That is what the scan tier exists to expose, and it was the largest instance on the disc: 206 banks' worth of samples found by hunting a magic while the 206-entry sector index in the entry's own head went unread. It is walked now, so read the lesson off the tier rather than off that entry - quote structural, or carry the scan figure beside the accounted share, because the two can differ by five points of the whole disc without the headline saying so. With that entry, the menu overlay's two pinned atlases and the standalone BGM sequence bound, no byte of the disc's accounted share now rests on a magic guess.

A low percentage that is the disc's own padding. The largest unwalked region on the whole disc was the monster archive, at well under two thirds accounted - and every unclaimed byte of it was zero, in one run per slot, ending exactly on a slot boundary. The battle loader seeks by a fixed stride and then reads a literal forty sectors, so a slot arrives whole however little of it the compressed block uses, and the file's extent is an exact multiple of that stride. The tail is declared slack, and the two battle side-band streaming files are the same shape at their own stride. Claiming it moves the structural share by more than five points of the whole disc and moves the worklist by nothing, which is the honest reading: the largest thing left was never a format.

A code image can be mostly fill, and a dump can agree with fill. nop encodes to a word of zeros, so a dump whose head is nop agrees with zero fill in any image at any base - and the dump corpus holds such dumps, taken over an image's own zero region. One confirmed a 20 KB extent inside a 131 KB hole in the cutscene overlay and carried most of that entry's reported code share. A confirmation now needs a non-zero word, and an all-zero extent is never claimed as code. A residue run is also cut wherever a sector or more of fill sits inside it, so the fill is reported as padding instead of being glued to the content beside it and counted as work.

An ascii_text run is not necessarily a string pool. The shape test counts NUL as printable, because a string pool is short strings separated by terminators - but so is a mostly-empty region with a few strings in it, and the all-zero test will not take it because one non-zero byte disqualifies the run. Read such a run by its zero fraction before treating it as unwalked format, and do not widen the all-zero test to absorb it: that moves bytes out of the worklist by redefining the instrument rather than by understanding them.

The sweep also corroborates a format page from the other direction: every pochi filler slot is exactly one 2048-byte sector, and the fill classifies as ascii_text while repeated_fill is tested first - so the sector is text-shaped rather than one short pattern repeated. A second, independent reason none of those slots carries a parseable asset.

Three families that read as unwalked format and were not

Each of these headed the ranking at some point and none of them was an unrecognised format. All three were bindings: the bytes were already understood somewhere and nothing connected that understanding to the walker. That is the shape to expect at the top of this ranking.

FamilyWhat the bytes areWhat was missing
18 lzs_container entriesthe ordinary scene-bundle descriptor table, at descriptor counts the detector's window excludes - retail bounds the count nowherea walker bound to the class
266 filler slotsone 1927-byte dev fill file plus 121 bytes of the mastering buffer's previous contentsa shape rule - the fill's line is 52 bytes long, so the period-1/2/4/8/16 test never fires and 266 sectors of filler ranked as work
2 headerless stillsfour 320x64 16bpp band uploads to VRAM (384, 0)a walker - the rectangle was already recovered from the consumer's immediates
the multi-bank VAB206 banks on a sector index the entry's own head carries, each one a two-chunk streama walker - the detector already read the count and the sector table, it just never claimed anything with them
the boot logo packfour publisher-logo images at offsets the parser already knewnothing about the format: the entry is a runtime overlay and a container, and the overlay override silently replaced the container's walker whenever a dump directory was supplied - which is exactly when the sweep runs
the menu overlay's two atlasesthe save-menu UI atlas and the save-slot icon sheet, both at named constants in this workspacea binding again: the dump corpus is the parser for a code image's code, so anything in its data segment fell to the magic sweep and read as "found by guessing" for assets already pinned. When the sweep finds something, check whether a constant for it exists.

An overlay's uninitialised data region travels with its code

The largest zero runs in the overlay images are the disc's own .bss, and the loader is what puts them there: it asks for the entry's sector count - the gap to the next entry, the same expression the TOC reader uses - and hands that straight to the transfer. So a linked image's uninitialised data rides into RAM with its code as zero fill. Those bytes are not a format nobody has walked; they are the buffers the image writes at runtime.

Shape cannot establish that, because zero fill looks the same whoever wrote it, so the claim rests on the image's own code. A claim is exactly one maximal all-zero run of at least 256 bytes - one non-zero byte splits it in two - and the image must address that run from two distinct addresses, or one address formed at four separate sites. A lone lui pair landing somewhere in a multi-kilobyte window is a coincidence an image with thousands of pairs produces. The pair scan walks forward from each lui, because the STR overlay hands its VLC unpacker the destination in a jal delay slot, which a backward-only scan never sees.

The worked case is the disc's largest single unclaimed run: 131172 zero bytes in the STR/MDEC cutscene overlay, between its initialised data and the unpacker at the top of the image. Six regions sum to it exactly, and every boundary is an address the overlay forms: four bytes of alignment, the play loop's 0x50-byte decode context (the ctx the cutscene page tabulates), its two 0x7800 slice staging buffers, four globals, and the 0x11000-byte STRv2 VLC lookup table - which the entry's own second residue run, a 3597-byte compressed blob, is unpacked into at load time. Two other zero runs in the same entry are refused for carrying no formed address, which is the answer the rule is supposed to give.

A bundle's last sector is the packer's buffer

Every scene bundle carried exactly one residue run, up to about 2 KB, and together they were the largest class on the worklist. Each starts where the last descriptor's LZS stream stops being consumed and runs to the entry's end, inside its last sector; no consumer reads it. The bytes are an earlier entry's: the packer used one buffer for the whole of PROT.DAT, in TOC order, and never cleared it, so offset k of an entry's slack holds the byte the nearest earlier entry whose extent reaches k holds there. That prediction has no free parameter and reproduces every byte of the slack in all 90 bundles; the first bundle, with no earlier entry that long, has zero slack - the buffer as first allocated. The same test, asked of a mapped overlay's last sector, finds the battle overlay's last 1604 bytes (above the slot-B base) to be the side-band stream 0894's, a donor the overlay-only comparison cannot see.

The battle overlay's head is twenty-two jump tables

The battle overlay's first 3.5 KB used to read as nine jump tables, counted as runs of in-image code addresses - and adjacent tables with no padding between them merge into one run. Every consumer bounds its index with sltiu before a jr, so each table's extent is that immediate times four, read off the dispatch: twenty-two consumers, twenty-two bases, tiling the region. They are claimed from named constants that are re-derived from the image's own instructions, and every one of the 850 arms lands inside the dumped function that holds its own jr.

A string is claimed by the address its image forms

An overlay's string pool has no header and no count. What makes a byte run a string of this image is that its code computes the address with a lui pair, and what bounds it is its own NUL - the same pointer-forming test the uninitialised-data claim uses, plus one level of indirection for a pointer table of such strings. It closed the battle tutorial's prompt pool (its consumer is the module's own code after all), the menu overlay's option strings, and the name tables of two dev modules.

A record is sized by the call that receives it

Three callees fix the extent of the record they are handed: the spawn routine and its pool wrapper take a spawn record in $a2 (header plus a move-VM program, bounded by its terminator), and the actor allocator takes a 24-byte static template in $a0. Every mapped image now has those records claimed off the call itself - including pointers staged in a saved register and copied over (move a2,s3), pointers loaded in the delay slots of switch arms that jump to one shared call, and image-local wrappers that forward their own argument. The slot-B band's own walk now follows the same shapes (and a pair completed in the call's delay slot). A loop that counts i from zero to a constant over a formed base states its array's length the same way, off the slti bound.

A runtime index states its stride, and a consumer its count

Most retail data is indexed by runtime values, but the index arithmetic still states the element size: the operand summed with a formed base is evaluated backwards as k * i + c through shifts, adds and copies - the compiler's ((i << 3) - i) << 2 is i * 28 - and each access through the element address names a field. The count comes from a bound check on the same index when the consumer has one, and otherwise from the next address the image forms that is not one of the array's own fields, kept only when that closes a whole number of elements. That sizes the field overlay's 27-record window table and its 73 sound-test rows. A loop that bumps a pointer instead of scaling an index states its length too: the Baka Fighter image walks seventeen per-fighter blocks of nine 0x60-byte move records through a pointer table, the 14 KB an earlier reading left as a "sparse table nobody addresses".

Two more claims close records nothing points at directly. A run between spawn records is read as a chain of [header][program] records and claimed only when it lands exactly on a pointer-credited record or on padding - both ends pinned. And the menu and field overlays each carry a window-program interpreter; the programs they are handed, most of them from switch arms, are claimed off those calls.

The formed-address test under all of this follows the lui register through copies, one indexing addu and branches, and drops it at any other write, the walk the reference scans share. And a label-credited dump that opens on eight nops is no longer counted as code: two such dumps, in the fishing and Baka Fighter images, had credited about 22 KB of data as code, which is why those two entries' residue went up when the account got more honest. Nor is an extent whose words mostly write $zero or decode to no instruction at all - compilers emit neither.

Two more residue runs that were not formats

A dev module's roster. The OTHER3 dev module read as 10904 bytes of ascii_text in two runs, which invites "a text blob nobody claims" and is wrong about the shape: three quarters of the entry is a fixed-stride table, and the stride is in the drawing loop's index arithmetic rather than in the bytes - (i << 5) + i then << 2 is i * 0x84, and a reciprocal divide wraps the cursor mod 81, which is the record count. The labels are Japanese, stored as little-endian Shift-JIS code units, and that is not evidence of a foreign build: 34 of the entry's 37 distinct calls into the executable land on a real function head, where the one image that really is from another build scores 0 of 42. The dev modules were simply never localised.

Another image's code. The slot-machine overlay's 1760-byte plausible_mips run has no prologue, no jr ra and no caller, and runs to the last byte of the entry. None of the three readings that invites - a jump-table body, data, or the interior of a neighbouring function - is right: those bytes are the fishing overlay's, byte-identical at the same file offset, an inherited tail from a buffer the packer did not clear. Both instruments cut such a tail now, by the same measurement, and the claim carries the donor's own name - so a plausible_mips residue run that survives here is this entry's own bytes rather than a neighbour's. On the slot-B images the cut also feeds back into the walk: a spawn record at or above it belongs to the donor, and four band images report slightly more residue once the record chain stops there.

Ratcheting the figure

A sweep nobody compares is a number in a terminal. A checker reads the sweep's CSV and ratchets it against a committed baseline, the way the disc-coverage gate ratchets the code figure. It is the third denominator, and it exists because the other two cannot ask this question: the port catalog measures the addresses this project has cited, and disc coverage's data half measures format recognition - a claim an entry satisfies fully while most of its bytes have never been walked.

FigureDirectionWhy
structural shareup onlybytes a parser walked to, from a header or a table
accounted shareup onlystructural plus magic-sweep hits - always the larger number, and never the headline
non-slack residuedown onlyunconsumed runs whose shape is not padding

Every figure is re-weighted by bytes rather than averaged over entries, and each is carried per format class as well as whole-disc. The per-class row is what makes a regression attributable: one class losing its walker moves the whole-disc figure by a rounding error, and a falling total does not say which parser to open.

The gate consumes the sweep's CSV rather than running it, so it skips where the disc is not - and a CSV swept on an older tree reports that tree's parsers through a passing gate, which is why re-running the sweep is part of taking a baseline rather than a separate habit.

Tests

A fixed set of entries is accounted off extracted/PROT with a floor asserted on each one's accounted share, plus the invariants that make the figure mean anything: claims stay inside the buffer, residue plus accounted equals the size, and structural never exceeds accounted. It skips and passes without disc data, like every other disc-dependent test here. The range algebra and the residue classifier are unit-tested on synthetic buffers, including the ordering traps above.

See also