2026-10-03 · schema v5 (unchanged) · security fix · upgrade: dotnet tool update -g RDLL.excelcannon
- Security: a TLS dependency updated (RUSTSEC-2026-0285, low severity). When you read a workbook straight from an
https:// URL, from the CLI or the .NET SDK, the TLS library behind it (rustls) accepted some handshake messages unencrypted that it should have refused. The handshake is still authenticated, so this cannot be used to tamper with or hijack a connection, but it is a protocol violation and 0.13.1 closes it. Reading a workbook from a local file makes no TLS connection and was never affected.
- Nothing else changes. Every model, report and render is identical to 0.13.0 apart from the
engine=0.13.1 stamp.
- Releases now refuse to ship a known advisory. The publish pipeline runs the same supply-chain check as CI, so a release cannot go out while a dependency has an open advisory, which is how 0.13.0 shipped with this one.
- Native engine:
win-x64, linux-x64 (glibc 2.17+) — unchanged.
- Upgrade:
dotnet tool update -g RDLL.excelcannon · dotnet add package RDLL.excelcannon.sdk --version 0.13.1
2026-09-25 · schema v5 (unchanged) · finishes #23 · closes #40–#44, #46–#48 · upgrade: dotnet tool update -g RDLL.excelcannon
.xlsb workbooks now carry their formulas. A binary workbook stores every formula as a parsed token stream, and until this release a computed cell came out as its cached value with an fx: true marker and no text. The new decompiler turns the stream back into exactly the text Excel writes for the same formula in an .xlsx of the same save: cell formulas, shared and array formulas, defined-name targets, validation bounds, conditional-format conditions, structured references (Sales[[#Headers],[Price]]) and references into other workbooks ('[1]My Data'!B2).
- So
diff and lint now cover formulas on .xlsb too. The formula category is scored rather than marked not compared, and broken-refs, computed-overwritten and region-holes run instead of being named in rules-not-evaluated. Across 39 .xlsb/.xlsx twin pairs — one workbook saved twice by one Excel — 38 now diff to zero findings on every axis; the one left is eight custom-view names Excel re-mints on every save.
- Never a guessed formula. Every token shape it spells was measured against Excel's own spelling. A shape it has not measured is declined by name through the existing diagnostics (1023, 1024, 1026) rather than published as a plausible formula that might be wrong. No fixture and no corpus workbook declines anything.
- The cost: an
.xlsb now does the same downstream work on formula text an .xlsx always did. A 10,000-formula chain extracts in 152 ms (35 ms before, when there was no formula text to process; the .xlsx twin takes 191 ms); a 993,727-cell model with 21,070 formulas takes 1.96 s rather than 1.59 s.
migrate from .NET. ExcelCannon.Migrate(oldJson, newJson) and the matching IExcelCannon.Migrate overloads, with MigrateOptions (sheet and region scope) and MigrationResult (IsMigrated, per-status counts, SameGeneration, the report verbatim on Json). If you implement IExcelCannon yourself — a test double — it gains two members.
- Fixed: an inline string's phonetic reading was read as part of its text. A cell stored as an inline string with a furigana guide came out as
東京トウキョウ instead of 東京. Shared strings were always read correctly.
- The codebase now meets every standard it publishes. Every function under cyclomatic and cognitive complexity 21, every file under 500 lines, 100% line coverage of the engine and of its tests, zero duplicated blocks, zero dead code, zero .NET and PowerShell analyser findings, and zero surviving mutants across all 6,412 mutation sites: every deliberate small change to the code is now caught by a test, fails to compile, or makes the code hang. The gate that measures these is now absolute: any regression fails the build. The numbers are on the standards page.
- Native engine:
win-x64, linux-x64 (glibc 2.17+) — unchanged.
- Upgrade:
dotnet tool update -g RDLL.excelcannon · dotnet add package RDLL.excelcannon.sdk --version 0.13.0
2026-09-01 · schema v5 (unchanged) · closes #29, #39 · upgrade: dotnet tool update -g RDLL.excelcannon
- Every artefact now says which build produced it.
schema_version: 5 tells you how to read a document; it never told you what wrote it. Six releases have emitted schema_version: 5 while disagreeing about what they publish, so a stored extract, diff or lint report could not tell you which side of a change it was on. New engine_version appears on the model, on all three report types, and in the text render header: excelcannon v5 | engine=0.12.0 | kind=schema | sheets=4.
- Absent means one thing only: written before this existed. Every model you already hold still loads. And a render of a stored model names that model's engine, not the binary you happen to be running — the distinction that matters when you go back to old evidence.
- Upgrading never changes a score. Two models of one workbook produced by different builds still diff to 100% and exit 0 — the stamp is never a difference, and a regression test now pins that so a future refactor cannot quietly break every stored baseline. A version mismatch surfaces as a warning on the report instead, where it belongs.
- The stamp does not destabilise your snapshot suite, and that was tested rather than argued. With the version temporarily set to 9.9.9 the full suite passes and not one of the 51 committed snapshots moves; regressing the stamp to absent fails 27 tests. The version is hidden from snapshots; the field is not.
- Fixed:
header_source was missing from schema columns (reported against 0.9.0). It was being suppressed at its default — and since header_cell is the commonest value, on an ordinary workbook the key simply never appeared. That made an absent key ambiguous between "this name came from a real header cell" and "this model predates the field", which defeats the whole point: telling a real name from a fabricated one without resorting to a regex. Same fix for inferred_type on input fields, which had the same guard while the schema-column version did not — so one fact appeared or vanished depending on which surface you read.
- Upgrade:
dotnet tool update -g RDLL.excelcannon · dotnet add package RDLL.excelcannon.sdk --version 0.12.0
2026-09-01 · schema v5 (unchanged) · closes #27, #32, #34 · upgrade: dotnet tool update -g RDLL.excelcannon
- Labels are now found where templates actually put them. A field's name was taken from the cell beside it. On real insurer templates the meaning usually sits in a block header above the field, or in a merged section banner — which is why 93.6% of one template's writable fields had no usable name. The search now looks there too. Across our 123,317-field corpus that is +1,373 named fields, and it never regresses on any single workbook.
- Nothing you already had gets renamed. The wider searches run only where the existing ones found nothing, so every field that had a label keeps exactly the label it had. New
label_path carries the composed chain — ["Purchase Request", "Requester"] — in the JSON, so you can choose your own granularity. It is deliberately not printed in the text render: on a single-banner form every field would gain the same prefix, which is noise rather than information.
- A wider search that reaches too far is worse than one that misses, because it steals a neighbouring field's label and shows up as better coverage. The search stops at a structural break — an intervening input, or a 20-row blank band — and every newly-labelled field in a corpus sample was checked by hand against what the workbook actually says. That check found a real bug: a cleared cell was naming fields with the empty string.
--section surface: the render, minus everything a prompt does not use. The text render is what goes into a model prompt, so its size is an operating cost, and there was no way to ask for less of it. --section and --exclude-section now work on extract and render, and the surface group emits the writable contract — input surface, names, headers, schemas, validations — with no grid, no style dictionary, no column widths and no dependency summary. Both commands also report the render's size, so you can budget without measuring.
--sheet, --exclude-sheet and --region on extract. diff and lint have had them for releases; extract did not, which was the surprising way round — you could scope your comparison but not your extraction. The workbook is still read in full and pruned afterwards, so a dropdown backed by a lookup sheet you did not scope in still resolves. That ordering is the whole correctness of the feature and it fails silently if got wrong, so the test that proves it was written first.
extract --batch and lint --batch: one process, many workbooks. A manifest of paths or URLs, one aggregate report, and a workbook that cannot be read no longer stops the run — its entry records the failure and the rest continue (--fail-fast opts back in). Results are ordered by manifest position, never by completion.
- An honest note on batch. The ~90ms per-invocation floor that motivated this is mostly the dotnet-tool shim paying .NET startup before it reaches the engine — the native binary is single-digit milliseconds warm. If you can call the SDK in-process you already avoid all of it, and batch is the fix for the case that cannot: a script.
--jobs, --pairs, --report-dir and --timings are deliberately absent rather than accepted-and-ignored.
- Upgrade:
dotnet tool update -g RDLL.excelcannon · dotnet add package RDLL.excelcannon.sdk --version 0.11.0
2026-09-01 · schema v5 (unchanged) · closes #35, #38 · upgrade: dotnet tool update -g RDLL.excelcannon
- New verb:
migrate — what a stored mapping has to change between two generations of one template. diff tells you how alike two workbooks are. That is the wrong question when a template gets revised and you are holding a mapping built against the old one. migrate old.xlsx new.xlsx answers the right one: which fields moved (with both anchors), which vanished, which appeared, and whose permitted values changed. Arguments are old/new rather than left/right, because direction is load-bearing here.
- It reports on the writable surface, not on content. No values, no cell-level findings, no fill state — a template filled twice with different data has nothing to migrate, and says so. Run against 0.9.0's
template_fingerprint, the report names both generations in its header so you can see at a glance whether you are comparing what you meant to.
- A
col_G placeholder is no longer treated as a name when matching fields across generations. 0.9.0 made fallback labels detectable; this is the first thing to use that. Aligning two fields because they share the label col_G would be worse than not aligning them at all.
- Bounding a diff can never quietly shrink a migration report. Since 0.8.0 a diff report's findings may be a sample. A migration report built on one would under-report changed fields with no sign it had — so
migrate does not read findings at all, and there is a test that bounds a diff to nothing and asserts the migration report is byte-identical regardless.
- New:
--withhold-name, declared suppression for schema mode. Schema mode decides what to withhold by inferring what looks like personal data. If your no-member-data guarantee rests on that inference, it rests on a guess. You can now name the columns and fields that must be withheld regardless — matched exact and case-fold only, after trimming. Member DOB matches member dob; it does not match DOB, Member D.O.B. or MemberDOB. There is no fuzzy matching of any kind, deliberately.
- The header stays, the data goes. A consumer needs to know the column exists, so its name is kept while its distinct set and its resolved vocabulary are withheld — on the schema column and on any input field the same validation backs. The render marks it
values(withheld:declared), distinguishable from the cardinality and shape guards, and the model records which declaration matched what.
- A declaration that matches nothing tells you so (diagnostic 1027) and suppresses nothing — a name that has drifted out of date is a thing you want to hear about, not a silent no-op.
- Where declared suppression stops is written down. A declaration names a header or a label, so it cannot reach a defined-name lookup list or a validation prompt — both pre-existing gaps of schema mode. A validation spanning several columns keeps its vocabulary unless every one of its ranges falls inside declared columns. And an inline
list("a,b") is left alone, because its members are already spelled out verbatim on the same line. A suppression you rely on has to say where it ends.
- Upgrade:
dotnet tool update -g RDLL.excelcannon · dotnet add package RDLL.excelcannon.sdk --version 0.10.0
2026-08-31 · schema v5 (unchanged) · closes #28 · upgrade: dotnet tool update -g RDLL.excelcannon
- The model now identifies the template, not just the file.
source.sha256 changes the moment anybody types a value, so it could never answer "which generation of this template am I looking at" — the question you ask when a filled workbook arrives and you need to know whether a stored mapping still applies. New source.template_fingerprint is computed over addressable structure only: sheet names, order and visibility, used column spans, frozen panes, merges, validation kinds and ranges, table names and column spans, autofilter spans, defined names. Not widths, not styles, not values.
- So a template's blank and filled workbooks fingerprint identically, while two generations of it do not. Both properties are corpus tests, not claims.
source.template_fingerprint_inputs ships beside the hash naming every contributing input, so you can tell in advance whether a cosmetic edit will invalidate a stored mapping.
- Two honest caveats, stated rather than buried. The
tf1: prefix versions the input set and how it is spelled — not the readers, so a future binary-format fix that starts reporting structure this build drops will move fingerprints under tf1: with no template change. And a fingerprint is necessary, not sufficient: two structurally identical templates share one, so a match is not proof of identity.
- You can now tell a real field name from a placeholder. A field labelled
col_G, and a schema column published with its column letter as its header, were indistinguishable from a field genuinely called col_G and a column a client genuinely named X — the only way to spot a fallback was a regex on the text, and a regex is a guess. New label_source and header_source say which it is.
header_source has three answers, because two would have lied. A column whose header cell is empty but whose name comes from the workbook's table part is not the same thing as one named after its own cell, and not the same as a fabricated column letter. Those are table_part, header_cell and column_letter respectively.
- Input fields now carry
inferred_type, so you can tell whether a field wants a date, a number or text — previously available on schema columns only.
- The render says whether a surface is addressable, not just how full it is.
= INPUT SURFACE 6 fields | 4 blank | 1 partial | 1 filled gains | 6 named, with letter-fallback and unnamed counts appearing when they are not zero. A rich template and an anonymous one used to look the same on that line.
- No label text changed anywhere. This release adds provenance; it does not move names. Verified field by field against the published 0.7.2 binary across 239 input fields in 20 corpus workbooks — zero differences. Widening where labels are found is separate work, and it is measurable precisely because this release shipped first.
- Stored models still parse. Every new field is optional on read and omitted at its default, so a model extracted by an older build loads unchanged and an unset value adds no key.
- Upgrade:
dotnet tool update -g RDLL.excelcannon · dotnet add package RDLL.excelcannon.sdk --version 0.9.0
2026-08-31 · schema v5 (unchanged) · closes #37 · upgrade: dotnet tool update -g RDLL.excelcannon
- You can now bound what
diff puts in its JSON, without touching what it scores. TopN only ever capped the Markdown; the JSON always carried every finding. On a benchmark that weighted format at zero, that meant 12,687 format findings — 80% of the payload — riding along in every stored artefact contributing nothing to the score. Two new controls: --finding-limit-per-category (SDK: FindingLimitPerCategory) and --exclude-finding-category (SDK: ExcludeFindingCategories).
- Scores never move. Every percentage is computed from the full comparison before anything is dropped — bounding is a reporting concern, not a scoring one. A bounded report and an unbounded one of the same pair carry identical numbers, and there is a test that fails if that ever stops being true.
- A bounded report still tells you what it dropped. A
truncation block states the true pre-truncation totals per category and per sheet, how many were retained, and why each was omitted (limit or excluded). You can always tell a report of 3 findings from a report of 64 that kept 3.
- Bounding cannot turn a real difference into a pass. This is the part worth reading twice.
is_identical and the exit code were both defined on the findings that survived — in the engine and in the SDK, where IsIdentical was literally findings.GetArrayLength() == 0. Excluding a category on a workbook whose differences were all in that category would have emptied the array and reported a clean, exit-0 pass on a workbook that genuinely differs. Both now read the full-comparison total. Excluding all four categories on a pair with 147 real differences retains nothing and still exits 1.
- A zero score weight still shows you the findings. Weighting a category to nothing means "do not count this towards the number", not "hide it". Suppression is opt-in, always explicit, and never inferred from the weights.
- Which findings survive is deterministic, and spread rather than truncated. A budget is filled round-robin across sheets before it fills any one of them, so a budget of 10 on a workbook with differences on eight sheets shows you eight sheets — not the first ten findings of the first sheet. Two runs of the same bounded diff are byte-identical.
- Off by default. An unconfigured
diff produces byte-identical output to 0.7.2 — no new key appears unless you ask for one. Measured on a real corpus pair: 40,257 bytes unbounded, 7,391 bytes at three findings per category.
- Upgrade:
dotnet tool update -g RDLL.excelcannon · dotnet add package RDLL.excelcannon.sdk --version 0.8.0
2026-08-31 · schema v5 (unchanged) · upgrade: dotnet tool update -g RDLL.excelcannon
- A dropdown that lists its choices inline now tells you what they are. A list validation can name a range —
list(Owners!$A$1:$A$6) — or it can spell the choices out in place: list("Buy-in,Buy-out"). ExcelCannon resolved the first into constraint_values (on an input field) and validation_values (on a schema column) and left the second empty, so the one case where the permitted values are written in plain sight was the one case you still had to parse out of the constraint string yourself. Both now publish their members.
- What changes in your output. On a workbook that uses inline list validations,
constraint_values and validation_values are populated where they were always empty. Nothing is renamed, no key is added or removed, and the schema version is unchanged at 5. If you hold a stored extract or schema baseline for such a workbook, expect a one-off diff on exactly those fields — the new reading is the correct one.
- What deliberately does not change.
Validation.values stays range-backed only, so no lint verdict moves — validation-violations judged inline lists correctly all along and still does. The text render is byte-identical: [list("Yes,No")] already states its choices, so the line does not grow a redundant = Yes,No after them. Exit codes are unchanged everywhere.
- It works in schema mode too, which is the half that could have gone missing. Schema mode withholds data values by design, and the code path that takes a withheld vocabulary back off a field asked a question that is permanently false for an inline list. Left alone, the fix would have worked in template and data mode and silently done nothing in schema mode. An inline list's members are declared in the validation rule itself — the same rule schema mode already publishes in full on the line above — so there was never anything to withhold, and there is now a corpus-wide audit that re-derives every published member from the rule beside it rather than taking the exemption on trust.
- Members are trimmed and empties dropped, matching the rule already used elsewhere:
list("N,,Y") publishes N and Y. A list that has nothing in it publishes nothing rather than an empty array, so "never looked" and "looked and found nothing" stay distinguishable.
- Upgrade:
dotnet tool update -g RDLL.excelcannon · dotnet add package RDLL.excelcannon.sdk --version 0.7.2
2026-08-12 · schema v5 (unchanged) · upgrade: dotnet tool update -g RDLL.excelcannon
- A binary workbook whose defined names could not all be read no longer makes
diff stop comparing formulas. 0.7.0 treated two different gaps as one. A .xlsb can hand over every formula in every cell and still contain one defined name — a union of areas, a link into a workbook that is not there — whose target this build does not spell yet. 0.7.0 read that as "the formulas are unreadable" and quietly dropped the formula score from the report, on a file whose formulas it had in full. If you diffed such a workbook, you got a report that declined to compare something it could have compared. That is fixed: the formula score is dropped only when the formulas themselves were not read.
lint now names one skipped rule where it named three. Same cause. On those workbooks the two rules that read cell formulas were reported as skipped when they had run perfectly well. The rule that genuinely cannot answer — broken-refs, which cannot tell an unread name target from a name pointing nowhere — is still reported as skipped, and --strict still exits 1, so nothing you gate a build on gets quieter.
- Two of the 130 workbooks we test against were affected, both real published files: an EIOPA taxonomy and an HM Treasury spending model. If you hold a stored
diff --json or lint --json for a .xlsb in that class, it will change — towards the answer the documentation already described.
- The
.xlsb acceptance gate grew to 35 twin pairs, 27 of them with zero differences. The new pair is one Excel had been refusing to open for a fixture-generator reason — a sheet name containing an apostrophe, written into a formula without doubling it. Fixing the generator got Excel to build the pair, and it passed on arrival.
2026-08-11 · schema v5 (unchanged) · closes #23
- ExcelCannon reads
.xlsb. Binary workbooks — the format Excel offers when a workbook gets large, and the one some regulators and insurers distribute their templates in — now go through extract, diff, lint and render exactly as an .xlsx does, into the same canonical model. No converting first, which matters because converting rewrites the number formats, validations and cached values you were trying to score.
- What comes across: cells and cached values, styles, merges, frozen panes, column widths and row heights, the worksheet autofilter, sheet visibility, defined names, data validations (with list vocabularies resolved through named ranges), conditional formats, table parts and VBA source. What does not, yet: formula text — the binary format stores a formula as parsed tokens rather than as the text you typed, and decompiling those is the next phase. A computed cell keeps its cached result and an
fx: true marker.
- Nothing pretends the gap is covered.
diff marks the formula category not compared rather than scoring it — the headline reads Overall match (formula not compared): 98.40%, so the number cannot be quoted without the caveat, and summary.not_compared says so in the JSON for a harness that would rather refuse. lint names the three rules that read formulas in its rules-not-evaluated notice and on skipped_rules, so lint book.xlsb --strict exits 1 instead of reporting Clean over checks it could not run.
- Proved against Excel, not against a spec. Every fixture is a twin: one workbook saved twice by one Excel session, once in each format. The acceptance gate is
diff a.xlsb a.xlsx over 35 such pairs — 15 built to exercise record shapes, 20 real workbooks including a 574-sheet EIOPA taxonomy — and it requires zero differences in values, formats and structure. 27 of the 35 come out with none at all; every difference on the rest is a formula the next phase will read, and the gate re-derives that rather than taking it on trust.
- Faster, not slower. There is no XML to parse, so the binary reader runs 1.3–3.8× quicker than the XML one on the same workbooks. A 5.4 MB real-world model with 993,727 cells extracts in 1.19 s.
- Fixed, and it affects
.xlsx too: a multi-line cell, a table column header with a line break in it, or a validation message spanning two lines was published with Excel's _x000D_-style escape left in the text. It is now decoded. If you keep an extracted baseline for a workbook with line breaks inside cells or table headers, expect a one-off diff — the new reading is the correct one.
- Upgrade:
dotnet tool update -g RDLL.excelcannon · dotnet add package RDLL.excelcannon.sdk --version 0.7.0
2026-08-11 · schema v5 · closes #24
- Linux is in. Both packages now ship a
linux-x64 native alongside win-x64 — the SDK's .so loads in-process in a Linux container, and the CLI tool runs under dotnet tool install on Linux (including the first-run executable-bit repair a Windows-packed nupkg needs). Cross-compiled with a glibc 2.17 floor (RHEL 7 / Amazon Linux 2 era — every glibc distro a .NET container plausibly runs; musl/Alpine and macOS remain out of scope for now).
- Byte-identical across platforms, proven: extract, diff and lint outputs on Linux byte-match the Windows binary's on real corpus workbooks — determinism holds with zero normalization, so baselines and reports are portable between OSes. The cross-build itself is bit-reproducible (a from-scratch rebuild reproduces identical binaries), and the glibc floor is asserted against the bytes actually packed, on every pack.
- The shipped-RID list is stated in the package descriptions, the README, this page and the SDK's
PlatformNotSupportedException message — all cross-checked in the packaging smoke so they cannot drift apart.
- Upgrade:
dotnet tool update -g RDLL.excelcannon · dotnet add package RDLL.excelcannon.sdk --version 0.6.2
2026-08-11 · schema v5 · closes #22
- A data cell is never a field's name. A member's postcode sitting left of a single-cell input was being promoted to that field's label — and labels travel: into the input surface, into diff JSON's
field_label, into everything a consumer treats as pure template structure. Two refusals now guard the label search: a cell in a column that a header above it names is that column's content (the search moves on and finds the real header), and identifier-shaped text — postcode, email, long mixed tokens — is an answer, not a question, wherever it sits.
- Refusals are stated, never silent — new diagnostic 1022 names the sheet, the cell and the rule that fired, and never quotes the text (a diagnostic carrying the postcode would republish what the refusal withholds). Fires in every mode; in schema mode strictly less is published than before.
- Narrowed by measurement, not worry: the blunt versions of these rules were tried against the 89-workbook corpus and rejected — one strips 559 genuine row labels, another 13,131 — and the rejections are pinned by tests. Known limitation, stated: a bare digit-run identifier (NI-number shape) in an unheaded label cell still publishes, because refusing digit runs strips genuine year-headings; real member extracts head their NI columns, which the structural rule covers.
- Genuine form sheets — label text left of an input, no table, no pane — label exactly as before, corpus-verified (88/89 workbooks byte-identical in template mode; the one mover is pure refusals).
- Upgrade:
dotnet tool update -g RDLL.excelcannon · dotnet add package RDLL.excelcannon.sdk --version 0.6.1
2026-08-10 · schema v5 · closes #21
- The same field is the same field, however long it is. Compare two extracts of one template — 455 members against 83 — and the input surface used to report every column twice: one field "removed", one "added", neither side credited with filling anything in. A field is now identified by where it starts, not by how far it runs, so that pair reads as one field both copies completed. On a 1354-field comparison with 28 such columns, the surface counts 1326 fields and "both sides filled this in" rises from 16.3% to about 18.8% — the same work, finally counted once.
- The size difference is still visible, because it is the interesting part. An aligned pair renders as one row with both extents on it —
4a. DEFERRED DATA!R9:R455 -> R9:R83 — and the JSON carries a_range/b_range whenever they differ. A pair both sides filled in completely but sized differently is kept in the table rather than filtered out as "complete", and ranks above the rows where the two sides simply agree.
- A missing side now says
null instead of saying nothing. A breakdown row used to omit the "a"/"b" key for a side whose workbook has no such field, which consumers flattened to "" — indistinguishable from "this side filled in nothing". Both keys are always present now. This is the breaking change in 0.6.0: reports written by earlier versions still parse, but a consumer that tested for key presence should test for null instead.
- A field whose start row moved (a row inserted above it) is re-matched by column, label and validation — but only when the match is unambiguous and the field is named. Two candidates, or no label, and the two stay two: a wrong merge would quietly rewrite the surface.
- Verified across the corpus: extraction and lint byte-identical on all 89 workbooks, and every diff report identical apart from the new explicit-null keys.
- Upgrade:
dotnet tool update -g RDLL.excelcannon · dotnet add package RDLL.excelcannon.sdk --version 0.6.0
2026-08-10 · schema v5 · closes reopened #20
- Pane-locked rows are furniture everywhere — no detected table required. The 0.5.1 clamp only acted where header detection had minted a body; on sheets where detection declines, lint still flagged the template's own locked annotation row and the input surface still built fields from it. Now the frozen pane itself is read as the author's statement:
validation-violations exempts pane-locked cells directly, and input fields clip to the first unlocked row — which also removes the phantom removed-and-added field pair that polluted diff tables when two extracts of one template differed in member count.
- Corpus-verified with a binary A/B of all 89 workbooks: lint byte-identical everywhere, 86/89 models byte-identical, and the three that changed are pure furniture removals (167 year-heading pseudo-fields in one ONS table, 80 field clips to unlocked rows).
- The residual field-identity question (same field, different data extent, still two rows) is now tracked as its own issue with a concrete design sketch.
- Upgrade:
dotnet tool update -g RDLL.excelcannon · dotnet add package RDLL.excelcannon.sdk --version 0.5.2
2026-08-10 · schema v5 · closes reopened #9, #11 and #20
- Header suggestions score what a row would actually publish (#9) — a units row of repeated
Pct/Auto cells can no longer outscore the real codes row above it: the scorer, the diagnostic message, and the schema's own publication rule now share one definition of "column-name-shaped". And the sparse-override warning (1020) counts effective names too, so the sparsest row in the workbook — the one the old suggestion pointed at — finally warns instead of being the one row that couldn't.
- Frozen panes clamp detected bodies (#20) — annotation rows locked above the pane can't land inside a table body anymore, so lint stops flagging a template's own units row as a violation. Verified across the corpus: 37 regions on EIOPA/EBA/statistical sheets moved their body to the first unlocked row, with everything ejected counted and stated, never silently dropped.
- Input-surface lines stop claiming date formats on text fields (#11) — the format-applicability rule now reaches the surface form templates actually render:
[list("Yes, No")] {mm-dd-yy} is gone, genuine date fields keep their formats, and the filtering happens in the model so SDK consumers get the clean JSON too. Corpus-wide: 1,520 false format claims withdrawn, zero added, zero rewritten.
- Upgrade:
dotnet tool update -g RDLL.excelcannon · dotnet add package RDLL.excelcannon.sdk --version 0.5.1
2026-08-10 · schema v5 · closes issues #6–#19
- Header detection got a second education (#9 #10 #17) — two-row header bands split by a spacer row now read correctly; diagnostic suggestions are scored against real content (never a blank row, never a row an override just rejected) and carry a machine-readable
suggested_header your pipeline can apply directly; a sheet with no anchor at all now says so (1019) instead of saying nothing; and a sparse override that names two of eight columns warns (1020) rather than silently under-describing.
- Column schemas you can trust (#11 #16) — no more
mixed_types on effectively single-typed columns, no date formats on string columns (including Excel's ;@ spelling), and dropdown vocabularies now appear on the column schema line itself. A lookup column typed at the top with a calculated fill below — the most common real-world shape — resolves as a literal_prefix vocabulary in every mode.
- Lint you can read (#12 #13 #18) — findings aggregate (one line for a 6-cell run of
#DIV/0!, not six), --top-n and --rule/--exclude-rule tame the floods, template annotation rows stop being flagged as violations, and linting a schema-mode model now states which value rules it could not evaluate instead of handing you a hollow exit 0.
- Diff sharpened (#14 #15 #19) — a scope naming a sheet absent from both sides refuses cleanly before any scoring; the per-field table lists differing fields first so
--top-n can no longer hide every real difference; and --weights value=0.5,format=0 reweights the headline from the CLI, with every report stamping the weighting it was scored with.
- SDK reaches everything (#6) — new
ExtractOptions (mode, header overrides, distinct values) makes the full extraction surface available in-process, byte-identical to the CLI. Platform honesty (#7): packages ship win-x64 today — Linux/macOS natives are roadmap, and every doc now says so. .xlsb (#8) is reported as unsupported-by-design (code 13) instead of "malformed".
- Upgrade:
dotnet tool update -g RDLL.excelcannon · dotnet add package RDLL.excelcannon.sdk --version 0.5.0
2026-08-09 · schema v5 · closes issues #1–#5
- Headers don't have to be on row 1 anymore. Title bands, summary rows and spacer rows above a table no longer defeat detection — frozen panes and filter ranges are read as the author's own statement of where the header sits, always corroborated by a style or type break so a wrong anchor can never publish data. When detection still declines but an anchor was visible, a new diagnostic tells you the exact
--header "SHEET!CELL" override to state it yourself. On real statistical workbooks this turned multi-thousand-cell suppression stubs into full column schemas. (#5)
- Dropdowns backed by named ranges now state their answers. A list validation pointing at a (often hidden) lookup range resolves to its literal values — in schema mode too, where the lookup sheet itself stays fully suppressed. Pure-literal sources only, capped at 256 with explicit truncation; ranges containing formulas, or overlapping described data, refuse to resolve.
lint now flags values outside these vocabularies as errors. (#1)
- Opt-in category vocabularies —
--distinct-values publishes the value set of low-cardinality columns (the "which four membership categories" question) behind cardinality, repetition and identifier-shape guards; off by default, byte-identical output when off. (#4)
- Score only what you're testing —
diff/lint gain --sheet, --exclude-sheet and --region; the scope is stamped into every report so a scoped percentage can never pass as a whole-workbook score, and a scoped sheet missing from one side is a finding, not a silent skip. (#2)
- Findings name the field, not just the cell —
field_label/field_range on diff findings, a per-field completion breakdown, and a constraint-violation flag when a produced value isn't in the field's (resolved) vocabulary — "wrong value" and "not a permitted value" finally separable. (#3)
- Hardening from the release audit: stale
customSheetView filters and panes are ignored (a saved view can no longer mint headers from data rows), and validation vocabularies never republish suppressed row values or hidden-sheet contents.
- Upgrade:
dotnet tool update -g RDLL.excelcannon · dotnet add package RDLL.excelcannon.sdk --version 0.4.0
2026-08-09 · schema v5
- Extract straight from a URL.
excelcannon extract https://…/book.xlsx?sig=… — presigned Azure Blob / S3 links work as-is (auth rides in the query string; no cloud SDKs involved). Works on diff and lint too, and in the SDK. Query strings and userinfo are never written to any output — models, reports, or error messages — so a leaked log can't leak the credential. 200 MB / 60 s / 10-redirect bounds; plain http:// refused unless you pass --allow-insecure-http.
- Errors your code can branch on. Failures now carry a stable category and code —
NotFound, AccessDenied, CorruptContainer, EncryptedWorkbook (password-protected files are finally told apart from corrupt ones), MalformedSpreadsheetML with the failing part/sheet/cell, and more. Codes are append-only forever. In the SDK: catch (ExcelCannonException e) when (e.Category == …) plus IsRetryable/IsUserFixable; on the CLI: --error-format json. The full table with "how to react" guidance ships in docs/errors.md.
- Extraction tells you what it glossed over. A new
diagnostics list (own 1000+ code range) reports non-fatal leniencies — files with no styles part, positional-only cell addressing, comments not modeled, skipped extension lists — aggregated with counts, deterministic, absent entirely on clean files. HasWarnings in the SDK, a one-line stderr summary on the CLI, and diffs ignore it.
- Hardened by a lint wall. The codebase now builds under clippy pedantic/nursery with unsafe code forbidden outside the FFI crate, unused-dependency and supply-chain checks in CI. 767 tests.
- Breaking: error code 4 is retired (each workbook failure now has its own code) and the Rust error type is restructured. The JSON model is unchanged (
schema_version 5) apart from the optional diagnostics and source.origin additions.
- Upgrade:
dotnet tool update -g RDLL.excelcannon · dotnet add package RDLL.excelcannon.sdk --version 0.3.0
2026-08-09 · schema v5
- Schema mode — share a workbook's shape without its contents.
excelcannon extract book.xlsx --mode schema keeps headers, formulas, validations, tables and geometry, and replaces every data-region value with a per-column description: type, number format, constraint, the R1C1 the column computes itself with, and how many rows were filled. For handing a workbook to an LLM, a schema reviewer or a ticket when the contents are personal or confidential. --data still works as an alias for --mode data.
- Suppression is deny-by-default, and says what it withheld. Text is published only for an affirmative reason — it names a column of a region that was actually described, or it clears a set of conservative caption checks — so a region the tool cannot read is withheld rather than printed. Hidden sheets are withheld whole. A
-- suppressed block counts every cell that went and what kinds they were, so nothing is dropped in silence. Audited against the full 89-workbook corpus: on the EIOPA DPM templates the schema render is 562 KB where the data render is 6.35 MB. It reduces exposure; it is not a certified anonymizer — header and caption text is published by design, and detection is heuristic. Read the output before sharing it.
- New package:
RDLL.excelcannon.sdk — the same engine in your process instead of as a child process. dotnet add package RDLL.excelcannon.sdk, then ExcelCannon.Extract / Render / Diff / Lint. The canonical JSON stays the contract, so nothing is duplicated as C# types. netstandard2.0 (.NET Framework 4.6.1+, Mono, Unity) and net8.0; the native engine ships in the package and failures arrive as ExcelCannonException rather than a crash.
- Dependency injection built in —
services.AddExcelCannon(o => { o.ExtractMode = ExtractMode.Schema; o.LintStrict = true; }) registers IExcelCannon as a singleton. Idempotent, your own registration wins, options validated at registration, and the interface is substitutable in a test host. No separate extensions package.
- Breaking: canonical JSON
SCHEMA_VERSION is now 5. template and data models are otherwise byte-identical to 0.1.0 — the additions (schemas, suppressed) only ever appear in schema mode. Re-extract stored models; text↔JSON round-trip remains byte-exact.
- Upgrade:
dotnet tool update -g RDLL.excelcannon
2026-08-08 · schema v4
- Legacy indexed colors now resolve to real RGB — regulatory workbooks (EBA COREP, older Excel files) use the legacy palette, and color often encodes meaning like cell editability; the model now shows
#RRGGBB everywhere.
- ~23% smaller text renders on style-heavy workbooks — style legends are compressed (per-sheet aliases, style families, inline props for rare styles), so more workbook fits in an LLM context window. On the EBA COREP annotated layout: 2.89 MB → 2.24 MB, one sheet's legend from 62 entries to 9.
- Breaking: canonical JSON
SCHEMA_VERSION is now 4 (style ids change because color resolution tightens deduplication). Re-extract stored models; text↔JSON round-trip remains byte-exact.
- Upgrade:
dotnet tool update -g RDLL.excelcannon
2026-08-08 · first release
- Extract a canonical, deterministic model from
.xlsx/.xlsm: cells, formulas, styles, merged headers, tables, dropdowns/validations, conditional formatting, named ranges, frozen panes, VBA macro source. JSON for machines + a compact text render for LLM inference.
- Understand how the workbook functions: detected column headers, cell roles (input / computed / label), a first-class input surface (every fillable field with its label and constraint), and dependency summaries.
- Diff two workbooks — or a blank template against a filled one — with category percentages (values / formulas / structure / format) and an input-surface completion score. CI-friendly exit codes.
- Lint filled templates: broken references, validation violations, formulas overwritten with literals, error values, blank inputs.
- Battle-tested against an 89-workbook corpus including EIOPA Solvency II QRT sets (574 sheets), EBA COREP/Pillar 3 layouts, and real filled bank/insurer disclosures.