Repository navigation
fix(373): adopter install + output fidelity fixes - #375
Draft
davidmatousek wants to merge 63 commits into
Draft
davidmatousek wants to merge 63 commits into
davidmatousek wants to merge 63 commits into
Conversation
Carries the define-stage artifacts onto the feature branch: PRD-373 (Approved; PM APPROVED, Architect + Team-Lead APPROVED_WITH_CONCERNS), the team-lead feasibility check, and the PRD index/backlog updates. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Spec for bundle #373 (K1-K3, K9-K15, #370), researched across four legs. Research corrected three PRD premises: the configured image models are retired or retiring, no code reads the templates' Gemini config blocks, and BSD cp can't write through nested links. Each is resolved within the PRD's goals by a spec ruling (S-1..S-13). The PM review's required changes RC-1..RC-8 are folded in. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…RNS) Plan, Phase 0 decisions (PD-1..PD-20), data model, four contracts (installer CLI, manifest completeness, extraction data, Gemini request and prompt scaffold) and a verification quickstart for bundle #373. Architect iteration 2. Rev. 0 was CHANGES_REQUESTED (2 HIGH / 8 MEDIUM); the fresh re-review of rev. 1 verified 25/25 resolved and folded N1: source-tree containment by file identity. PM rulings P-10.1..P-10.4 and the resulting spec touches are applied to spec.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
tasks.md (40 tasks, bucket-scoped by lane and carve unit) and agent-assignments.md for bundle #373. PM, Architect and Team-Lead are all APPROVED_WITH_CONCERNS. Architect iteration 3: F1-F4 and NM-1/NM-2 folded; K12 counts the normalized Section 7 map per FR-K12.1/K12.2. Team-lead re-cost: 3.5 / 5.2 / 7.7 d. TW-0 not fired (5.21 d vs 5.5 d), re-checked at W1 exit (T039, PM ruling P-11.3). OQ-5 closed. The spec, plan, data-model and contract touches from the tasks reviews are applied (P-10.4, P-11.1..P-11.3, AR-1..AR-3). /aod.analyze PLAN-exit gate: clean (CRITICAL=0 HIGH=0 MEDIUM=0 LOW=7). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
16 synthetic run-directory scenarios under tests/scripts/fixtures/fidelity_373/ covering K9 short-form headers and empty bands, K10 ### MAESTRO heading, K12 4c/4b baselines, bracketed statuses and the Section 7/tier ID mismatch, K13.1 partial and drifted Section 4, K11 funnel shapes (STEP-bound, strong-reduction, 3-tier, threats-only, volumes-unavailable, inherent-less join), status/score edge cases and an mmdc-free posture run. README.md records the hand-computed expected values for each scenario. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Stdlib-only symlink sandbox builders for the K3 installer tests: links at and above entries, nested, dangling, looping and wrong-type links, links into the source clone and its parent, case-variant links, a vendored clone and a dangling deprecated-command link. Symlinks are built at test time only. Kept in its own commit so a TW-7 early PR can cherry-pick it (LOW-3); Lane A' extends this file with the harness in W1. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
T001: literal totals for the modules about to be gated (all green: the four fast-workflow modules 75/75, the exec-arch payload module 12/12, tachi-pytest.yml's 16 modules 181 passed/1 skipped/1 xfailed) and the N4 ungated set (22 pre-existing reds, none F-373). Oracle pre-snapshot run list: 84/84 runs, deterministic across two runs; snapshot.sh appended. T002: eight-call W0 smoke, all HTTP 200 with an image on first attempt. IMAGE_SIZE_RESTORED = true; 3:4 at default size passes on both GA models; no non-404/403 status for AR-2; camelCase inlineData/mimeType. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…est guard (A-1) K1: INSTALL_MANIFEST.md gains the three skills that were shipping code but missing from the manifest (tachi-output-integrity/, tachi-misinformation/, tachi-human-trust-exploitation/), bringing coverage to 21/21 skills and agents. K2: the manifest gains scripts/populate-affected-assets.py, the fourth distributable script, with the dependency note corrected to "stdlib-only at import; PyYAML lazily for the PDF coverage-attestation page" and the "every other file in scripts/ is tachi-internal" rule spelled out. README.md and docs/guides/DEVELOPER_GUIDE_TACHI.md carry the same comment-free manual-install loop between BEGIN/END MANUAL INSTALL LOOP markers, so a `#` inside the block can never break a paste into an interactive zsh session (zsh does not treat `#` as a comment by default). tests/scripts/test_install_manifest_completeness.py is the fail-closed completeness guard (FR-K2.2-K2.5, S-13): every skill, command, agent and distributable script must be covered by the manifest or explicitly excluded with a reason; the manual-install loop must be byte-identical across both docs and comment-free under bash and zsh; a manifest-driven end-to-end install into an empty project must produce every required path. The negative matrix proves each assertion fails closed. .github/workflows/tachi-install-fidelity.yml gates this module alone (the manifest-completeness job) on pull_request and push:[main] through one paths anchor — this is the W1 cut line (PD-9, PD-20): nothing else may redden this commit's run (US-1, US-5). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…mp-copy fixture (A-2) Adds the extraction-fidelity job to tachi-install-fidelity.yml, gating test_tachi_parsers.py, test_extract_infographic_data.py, test_extract_report_data.py, test_extractor_contract_fixes.py and test_executive_architecture_payload.py on bare ubuntu for the first time (the R-3 safety net). This lands before any Lane B1 parser commit, per the W1 launch rule. test_extract_report_data.py's agentic_app_report_typst fixture (PD-8) now skips its five image-flag cases when mmdc is absent, and runs the extractor against a tmp_path_factory copy of the whole agentic-app sample-report rather than the tracked example in place — local runs no longer re-render tracked attack-tree/attack-chain PNGs (#365-adjacent hygiene, not a #365 fix). paths: gains examples/**, the exec_arch and golden fixture trees (N3), report_data/** and fidelity_373/** (C-1), schemas/taxonomy/*.yaml (OQ-5), and the five gated modules themselves, in lock-step with the new job (F-250). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…record (T004-T006) A-1 70a9e7a: manifest-completeness green on #375 (run 36372006524). A-2 2024ce4: extraction-fidelity green on bare ubuntu at first run (run 36373590935); no quarantine needed. The one-time manual-install loop record under /bin/bash 3.2.57 and zsh 5.9 is added (T005). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Move the Gemini image-generation request surface from the shut-down preview/GA chain to the current GA replacement pair per plan.md PD-14 and spec.md FR-K14.1/K14.4/K14.5/K14.6: - .claude/skills/tachi-infographics/references/gemini-prompt-construction.md: known-good request body (aspectRatio/imageSize nested under generationConfig.imageConfig, imageSize present per IMAGE_SIZE_RESTORED=true per the W0 live smoke, T002); key -> field table; chain + 404/403 walk set with default_model == chain[0]; camelCase inlineData/mimeType response parsing (SDK spelling accepted); a new Error Guidance summary; the corrected executive-architecture no-scaffold routing (never falls back to this file's own dashboard prompt); live-verification provenance and a single "Retired models:" line; resolution removed everywhere. - .claude/agents/tachi/threat-infographic.md: a Skill References row pointing to executive-architecture.md; a new "Request Configuration Mapping" instruction (read the active template's config, map it through the reference's table); the Error Handling & Graceful Degradation table gains 400 (not walked), 404/403 (walked), chain exhausted, and catch-all rows, plus the 429 reorder-models hint; two stale per-model JPEG/PNG claims naming retired IDs are generalized. - adapters/claude-code/agents/references/infographic-gemini-api.md (FR-K14.6): same body shape, chain and response keys as the reference. Static contract test coverage (A1, A4, A6, A7) lands with the T014 module in a separate commit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Update the six `## Gemini API Configuration` blocks to the GA chain
(model/fallback_model), per plan.md PD-14 and spec.md FR-K14.2:
- templates/tachi/infographics/infographic-{baseball-card,
maestro-heatmap,maestro-stack,risk-funnel,system-architecture}.md:
model -> gemini-3-pro-image, fallback_model -> gemini-3.1-flash-image
(response_modalities, aspect_ratio: "16:9" and image_size: "2K"
already matched the contract and are unchanged).
- .claude/skills/tachi-infographics/references/executive-architecture.md:
new `## Gemini API Configuration` section placed after the END
VERBATIM PROMPT BLOCK marker and before `## Payload schema` (plan.md
PD-1, outside the FR-212-6 lock) -- same chain, aspect_ratio: "3:4"
(portrait), image_size: "2K" per IMAGE_SIZE_RESTORED (T002 W0 smoke).
AR-2's decision 2 (3:4 default-size check) PASSED on both models
(test-results/w0-smoke.md), so this configuration commits now per
T013's contingency note.
- templates/tachi/infographics/INFOGRAPHIC_TEMPLATES.md:127-134:
reworded to say the static contract test validates the required
template sections, not the agent at runtime.
No prompt text is touched (K15 prompt-text edits are W2, after T007's
splitter hardening lands) and no fence was added under a prompt heading
(FR-K15.3).
Static contract test coverage (A2, A3, A5) lands with the T014 module
in a separate commit.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Anchor extract_prompt_scaffold's primary "DATA CONTENT" marker search to
line start (^DATA CONTENT \(render this) instead of an unanchored
substring find, so the IMPORTANT note's own "...DATA CONTENT sections."
prose reference can never be mistaken for the real marker.
Restrict the FOOTER search to start after the marker line (find(...,
marker_line_end) instead of a whole-prompt find), and drop the no-newline
find("FOOTER") fallback entirely, so a "FOOTER" that happens to appear
earlier in the preamble (for example inside a hard-wrapped K15
layout-label sentence) can no longer produce a false split.
Keep the bare ^DATA CONTENT fallback (N8, decided per L14): it is
unreachable while the primary marker matches on all five shipped
templates, retained only against a future preamble wording change. Pinned
in a code comment as required.
Verified byte-identical scaffold output on all five templates via the
existing golden byte-gate test (test_extract_infographic_data.py::
test_existing_templates_unchanged, part of the 37/37 file and the wider
87/87 gated-module baseline).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
K9 (FR-K9.1/K9.2, data-model.md §3): - New shared helpers in tachi_parsers.py: parse_score (Decimal | None, rejects non-finite NaN/Infinity), normalize_header + HEADER_ALIASES (the four canonical Coverage Matrix fields), _resolve_row_fields, is_placeholder_id, match_heading, _heading_level. - parse_markdown_table gains an optional start_line argument and a level-aware stop rule: scanning for a table stops at the next heading of the same or higher level than the matched line. A non-heading match (a handful of callers match bold paragraph text) keeps today's rule unchanged. Because a level-1-or-2 match reduces mathematically to the old hardcoded "## "/"# " rule, every pre-existing caller is unaffected by construction — pinned as regression tests for the three bold-paragraph callers, a representative "##" caller, and both scripts/generate-risk-scores-sarif.py call sites (FR-K9.2). - parse_compensating_controls_md's Coverage Matrix loop: skips placeholder Threat IDs, resolves Residual Score/Severity through the alias table, and its nested _score_to_band now delegates to parse_score. K12 (FR-K12.1-K12.5, data-model.md §6): - normalize_delta_status (N6 order: strip backtick/emphasis/whitespace, one [...] pair, the run again, upper-case), delta_status_by_id (the normalized Section 7 status map, has_status_column, row_count), apply_delta_status (badge/top_findings stamping only, no part in the counts — NM-1), warn_delta_scope (PD-16's scoped, aggregated warnings). - compute_delta_counts's signature changes to (status_by_id, resolved): it counts the normalized Section 7 map, never a tier's findings list (which has no delta_status key at all on Tier 1/2, and an unnormalized one on Tier 3). Both extractors' existing call sites are updated to build the map via delta_status_by_id first, so the gated modules stay green between waves; apply_delta_status/warn_delta_scope wiring is W2 (T017). - parse_resolved_findings now accepts both the current "## 4c." heading and the legacy "## 4b." spelling via match_heading, and skips placeholder rows. K10 (FR-K10.1): both extractors' MAESTRO layer distribution parsers now match the Section 6 heading at level 3 or level 4 via match_heading, instead of an exact "####" substring. Verified against tests/scripts/fixtures/fidelity_373/ (T003): exactly 2 findings (not 7 phantom rows) on the short-form/empty-band controls fixture; exact NEW/UPDATED/UNCHANGED/resolved counts on both the "## 4c." and legacy "## 4b." baseline fixtures, including every bracketed Status spelling; the ID-set-mismatch warning fires but counts stay unrestricted; a Status-less non-baseline Section 7 warns never; a "###" MAESTRO heading parses on both extractors. 150/150 gated-module tests pass (test_tachi_parsers.py 35, plus the existing 87 extraction-fidelity + 28 co-fired-gate tests), 2 pre-existing skips unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
parse_compensating_controls_md gains composites_by_id=None (data-model.md
§3): a {threat_id: Decimal} map of risk-scores composite scores used to
fill a row's inherent score by ID join when the Coverage Matrix has no
Inherent Score/Inherent column, or the cell is unparseable.
New classify_control_status (whole-token, case-insensitive; a "partial"
prefix wins first, then the recognized no-control set silently, then
"found" with no negation token, else "none" with a warning). It replaces
both copies of the row idiom: the STRIDE coverage matrix derivation and
the Section 1 coverage-summary fallback. The Section 1 table reader keeps
its own matching (it must skip rows that aren't statuses).
The Coverage Matrix row loop now: resolves Inherent Score/Inherent and
Control Status through the alias table (already defined in T016);
attempts the composites join when the column is absent or unparseable;
clamps residual to at most inherent once, with a warning; defaults a
missing residual to the inherent score (no credit) when inherent is
present, with a warning; bands both scores from the (possibly joined or
clamped) Decimal, falling back to the Inherent Severity/Residual Severity
columns only when the score itself doesn't parse. The legacy raw-score-
vs-heading "misclassified" check is removed — banding is now uniformly
score-derived, which is what the clamp-then-band rule supersedes it with.
Both tier-1 call sites now build composites_by_id from risk-scores.md and
pass it through: extract-infographic-data.py's extract_severity (which
already received rs_content unconditionally) and extract-report-data.py's
main(), which now reads risk-scores.md at tier 1 too (previously tier 2
only).
Each finding dict gains status_class, inherent (Decimal | None) and
inherent_severity; residual_score/residual_severity are unchanged in
shape but now reflect the post-clamp value.
Verified against tests/scripts/fixtures/fidelity_373/: the inherent-less
join path (funnel_join_inherent_less), the short-form fixture's inherent/
status aliasing, and the ten-row "kitchen sink" fixture's full
classify_control_status matrix, unparseable-score/missing-residual/
missing-inherent/clamp warnings (all five, aggregated and counted exactly
per the fixtures README's hand computation).
This is a stand-alone K11 commit (no other K-item shares it): T039 can
revert it cleanly if the W1-exit trip-wire carves K11, leaving T016's K9/
K10/K12 work and T024's K13-posture (built on top of K11's counts when it
ships) unaffected.
157/157 gated-module tests pass (test_tachi_parsers.py 42, plus the
existing 87 extraction-fidelity + 28 co-fired-gate tests), 2 pre-existing
skips unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
New compute_risk_posture(counts) -> (level, label) in tachi_parsers.py (data-model.md §5, D-3): the level is the highest non-zero severity band present — critical, else high, else medium, else low — mapped 1:1 to its label (CRITICAL RISK / HIGH RISK / MODERATE RISK / LOW RISK). Zero findings, or an all-zero counts dict, give low/LOW RISK. The function is deliberately agnostic to which counts dict it receives: the caller picks residual severity (after K11's clamp, when a controls report exists), the inherent composite, or qualitative counts (W2, T025). Verified here to run correctly on K11's post-clamp counts via a real parse of tests/scripts/fixtures/fidelity_373/posture_mmdc_free/ (the highest non-zero band is High). Stand-alone commit per the carve rule (K13-posture is the third carve candidate, after K15 and K11): T039 can revert it independently if the W1-exit trip-wire carves K13-posture, leaving K9/K10/K12 (T016) and K11 (T020) unaffected. 160/160 gated-module tests pass (test_tachi_parsers.py 45, plus the existing 87 extraction-fidelity + 28 co-fired-gate tests), 2 pre-existing skips unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… T016, T020, T024) C1 (K14 content) d6e94e7, c688144: CI green incl. maestro-coverage. B1 replayed after A-2 green: 1660c10 (PD-6), 1e30e90 (K9/K10/K12), a837ae8 (K11, stand-alone), 5328b56 (K13-posture, stand-alone); CI green incl. catalog-drift, mmdc-preflight and maestro-coverage. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… (T008-T010 lock-step)
D-1: install.sh now denies by default when a destination it would write
to is a symlink, or lies under one. Nothing is written until the
pre-flight is satisfied. --follow-symlinks opts back in, naming every
resolved destination before copying and again in the summary, and only
ever copies through a followed link -- it never deletes through one.
Even with the flag, install.sh always refuses:
- broken, looping or wrong-type links (a file where a folder is
needed, or the reverse);
- links nested inside an installed folder's destination subtree;
- destinations inside the tachi source clone (by identity, not path
string, so case variants and firmlinks cannot bypass it).
The version-checkout ref restore is no longer silent: the trap now
records the run's exit status, tries the restore, and -- if the
restore itself fails -- warns on stderr with the manual recovery
command, then exits with the original (not the restore's) status.
Test-first record:
specs/373-adopter-install-output-fidelity/test-results/k3-test-first.md
Suite: tests/scripts/test_install_sh_symlink_preflight.py +
tests/scripts/test_install_sh_ref_restore.py, 56/56 green on
/bin/bash 3.2.57 (macOS's bundled shell, the strict compatibility floor).
Implemented to the bash 3.2 constraints (no associative arrays, no
mapfile): resolve/phys_dest/under are only ever called inside a
conditional (a bare or `local`-masked assignment under set -e would
mask a dangling-link failure as a false "resolved"), the checked set
and every report are newline-delimited strings walked with
`while ... done < <(...)` rather than a pipeline (which would run the
loop in a lost subshell), and CDPATH is unset once up front so a
relative --source path can't leak a cd message into a captured
`pwd -P`.
CI: scripts/install.sh and its three test/helper modules are wired
into .github/workflows/tachi-pytest.yml's &hardening_paths and pytest
invocation in this same commit (F-250 lock-step requirement).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ity (A-3) Cherry-picked from Lane C1 (373-w1-laneC1 @ bbee14f) without its own commit: tests/scripts/test_gemini_request_contract.py implements A1-A8 per contracts/gemini-request-and-scaffold.md's static contract test table — the reference's known-good request body, the five template configuration blocks, the executive-architecture block's line index and aspect ratio, the adapter copy's conformance, the key-to-field table, the model chain, the retired-model scan, and the strengthened scaffold boundaries. No network call, no live key. Wires the module into the extraction-fidelity job's pytest invocation, last in W1's lock-step sequence (PD-9, PD-20). Extends the shared paths: anchor with the module itself and adapters/claude-code/** (A4 and A6 pin the shipped adapter copy). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-SESSION snapshot T039 (team-lead, P-11.3): TW-0 4.05 d vs 5.5 (margin 1.45 d), TW-1 T020 0.011 d vs 0.36, TW-2 T024 0.036 d vs 0.15: nothing fires, no revert. K15, K11 and K13-posture all ship in W2; only K15 stays carvable (TW-6). Wave-2 gated set at 0448bfb: 424 passed, 0 failed, 2 skipped (known), 0 regressions. T008, T009, T014 and T039 marked done; T010 converges in W2. NEXT-SESSION.md is the W1-exit fallback snapshot. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… run Run 36381945986 at 0448bfb: ubuntu-latest (bash 5, GNU coreutils) and macos-latest (/bin/bash 3.2, BSD) both success; no fix-forward needed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Add --follow-symlinks to the install section's usage/flags block, after --source and before the manual-install alternative. The scope paragraph is taken verbatim from contracts/installer-cli.md's Synopsis section (P-9.3 disclosure, P-11.1 wording), matching install.sh's usage() at the branch tip. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…hs glob-safely (SEC-K3-01, SEC-K3-02)
Addresses two findings from T011's advisory security review
(.aod/results/security-analyst-373-k3.md) against scripts/install.sh.
SEC-K3-01 (MEDIUM): resolve()'s 40-hop ceiling exceeded this platform's
real symlink-traversal limit. The review measured Darwin's actual
ELOOP threshold directly: `mkdir -p` through a chain succeeds at 32
hops and fails with "Too many levels of symbolic links" at 33. Because
the pre-flight's own ceiling (40) was more permissive than the real OS
constraint, a 33-40-hop `.claude` chain would pass classification as
`outside` (flag-eligible), and with --follow-symlinks the copy loop
would then hit the OS's real ELOOP mid-copy -- after already writing
an earlier, unrelated manifest entry (probe3_partial.sh confirmed a
non-atomic partial install).
Fix: lower resolve()'s ceiling from 40 to 32, the cross-platform-safe
minimum of Darwin's MAXSYMLINKS (32) and Linux's SYMLOOP_MAX (commonly
40). A chain the installer now classifies as resolvable is always one
the real OS can actually walk during copy, on either platform -- the
mismatch is closed structurally, not papered over. The contract's own
"40 hops" text (contracts/installer-cli.md, data-model.md) is intentionally
left untouched here per the orchestrator's routing: the architect ratifies
the new ceiling and amends that prose at the P0 checkpoint.
Added a regression test reproducing the exact failure mode: a `.claude`
destination reached through a 33-hop chain, run WITH --follow-symlinks,
now refused by the pre-flight as unresolvable before any copy is
attempted, proven via before/after tree snapshots on both the project
and the chain's target. This holds on both CI legs by construction --
the installer's own 32-hop ceiling refuses the chain regardless of
Linux's more permissive real ELOOP limit. Moved the existing
40/41-hop boundary test to 32/33 and updated its ids, docstring, and
the two related docstrings in install_sh_helpers.py.
SEC-K3-02 (LOW): strict_prefixes()'s `set -- $rel` under `local IFS=/`
expanded $rel unquoted in command-argument position, which bash
subjects to pathname expansion (globbing) in addition to the intended
field-splitting. A manifest entry whose path segment contained a glob
metacharacter (`*`, `?`, `[...]`) could silently glob-expand against
files in the target project (the installer's cwd when the pre-flight
runs) instead of being treated as a literal ancestor path segment.
Fix: rewrite strict_prefixes() to split via parameter expansion only
(`${rest%%/*}` / `${rest#*/}`), never `set --`, so pattern-matching
happens against the string value alone -- no pathname expansion is
possible regardless of what exists on disk. Verified behavior-preserving
against the original IFS-based implementation for both a normal
multi-segment path and a single-segment path (prints nothing).
Added a regression test with a manifest entry (`gl*b/thing.md`)
alongside a decoy symlink (`glob`, matching the glob pattern) pointing
outside the project -- exactly the shape the pre-flight would refuse
if wrongly classified. Confirmed the unpatched installer actually
falls for it (refuses citing the decoy: "glob -> .../outside-decoy
[outside project]"); the patched installer ignores the decoy entirely
and installs the literal entry normally.
SEC-K3-03 (the --version leading-dash guard) is out of scope for this
fix -- routed to the architect at P0 per the orchestrator's plan.
Verification: /bin/bash -n and shellcheck both clean; the three
x-release-please-version markers are byte-identical to 2b59e05;
tests/scripts/test_install_sh_symlink_preflight.py +
test_install_sh_ref_restore.py: 58 passed (56 baseline + 2 new);
test_install_manifest_completeness.py (e2e, runs install.sh): 10
passed. Both new tests independently confirmed to fail against the
pre-fix installer (git-stash round trip), proving genuine regression
coverage rather than vacuous assertions.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…018) FR-K9.3: on the explicit-path branch, /tachi.infographic's data-source detection now recognizes a compensating-controls.md by either the Residual Score or the short-form Residual column, matching the HEADER_ALIASES table K9 already applies in the two extractors (scripts/tachi_parsers.py). A short-form controls file passed explicitly no longer halts as UNABLE TO DETECT DATA SOURCE TYPE. Auto-detection by file name is unaffected. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
D-3 (data-model.md §5): every consumer of the risk posture now renders
metadata.risk_posture_level / metadata.risk_posture_label verbatim,
with color taken from the level alone.
- infographic-baseball-card.md: the badge shows the label alone
("{risk_posture_label}", e.g. "HIGH RISK") instead of the doubled
"RISK POSTURE: {risk_posture}"; the ">20% of findings" color rubric
is replaced by the level-only rule, in both the Zone Specification
and the Gemini Prompt Template's DATA CONTENT text (outside
prompt_scaffold).
- gemini-prompt-construction.md: the reference prompt gets the same
label-alone badge text.
- INFOGRAPHIC_TEMPLATES.md: the placeholder table points
{risk_posture_label} at metadata.risk_posture_label instead of the
stale "Spec Section 2: Risk Posture Indicator".
- infographic-specifications.md: the spec reference documents the two
new metadata fields alongside the existing risk_posture sentence row.
- threat-infographic.md: the agent's JSON contract example gains both
fields.
- schemas/infographic.yaml: both fields are added to the Metadata
section's required_fields.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
data-model.md §4.1-§4.3: the risk-funnel template and its skill reference now describe the real STEP/FLOOR width rule (STEP = 10, FLOOR = 30) instead of the old "10% minimum narrowing, 10% floor" formula, and key ghost-tier and unavailable-volume rendering on the `ghost` field / a `null` volume rather than "data source unavailable". - infographic-risk-funnel.md: the ASCII mockup and the Tier 2/3/4 zone headers drop the nominal ~75/50/30% widths (now dynamic); Tier 2's data source describes the one row set it shares with Tiers 3-4 in 4-tier mode; the width-calculation fence carries the real formula and the ghost cascade; the "0% risk reduction" note is documented as keyed on a numeric 0.0 (STEP still narrows tiers even at zero reduction — they are never flattened to equal widths); the sidebar's Risk Reduction line renders "not available" on `null`. - template-specific-formats.md: the baseball card's and the funnel table's Risk Reduction row/definition are now the row-derived Tier 2→4 volume reduction (S-9), not an Executive Summary total or an average of per-row deltas, "not available" on `null`; the Tier Width Calculation subsection and the "all findings same severity" and "zero risk reduction" edge cases match the same STEP/FLOOR rule and the ghost/0.0 keying. All edits sit under "## Zone Specifications" (before the "## Gemini Prompt Template" heading) or in template-specific-formats.md, which isn't one of the five scaffolded template files — prompt_scaffold is unaffected (verified below). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…w-on)
adapters/claude-code/agents/references/infographic-gemini-api.md is
the fallback prompt-construction reference distributed alongside the
agent (frontmatter: "extracted_from: Gemini API Prompt Construction +
Integration"), and Lane C2 owns it. It carried the same
"RISK POSTURE: {risk_posture}" doubled-badge text already fixed on
its source (gemini-prompt-construction.md) and on
infographic-baseball-card.md in the prior T026 commit (edd8f33). This
was missed there; caught during this stage's verification sweep for
stale references to the removed patterns.
Not one of the five scaffolded templates under
templates/tachi/infographics/, so prompt_scaffold is unaffected.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ded prompts (T028) Adds the final PD-17 wording (contracts/gemini-request-and-scaffold.md § K15 instruction text, rev. 1) as one new paragraph between IMPORTANT: and STYLING DIRECTIVES in each of the five scaffolded templates' Gemini Prompt Template preambles and in the reference prompt's Fallback Prompt Structure: - the layout-label sentence (do not render DATA CONTENT/FOOTER or any other instruction text as visible image text); - the allow-list rule (every finding ID and component name in the image must come from the ALLOWED IDS AND NAMES line; never invent one). Each sentence stays on one unwrapped line so no wrapped line can start with FOOTER (L14). No fence added, no early DATA CONTENT marker text introduced, no line starts with FOOTER outside the real footer. Verified with the real extract_prompt_scaffold(): all five templates still split found=true, postamble byte-identical, preamble gains only this paragraph. This intentionally drifts the five prompt_scaffold goldens in test_extract_infographic_data.py (T032 regenerates them at wave end). A1-A8 in test_gemini_request_contract.py stay green (52 passed), as do test_executive_architecture_payload.py and both maestro modules (23 passed, 2 pre-existing skips). Not in this commit: executive-architecture.md's PD-2 amendment, the lock-rule note, and the agent's executive-architecture section — T028 stage 2/3. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
tasks.md T016 requires parse_threats_findings (tier-3) to normalize its Status column at parse, same as delta_status_by_id's Section 7 map. It was storing the raw stripped Status instead, so tier-3 badges and the infographic's tier-3 top_findings[].delta_status showed raw "[NEW]"-style values instead of "NEW". Counts were unaffected (they come from the normalized Section 7 map, not this field). Fix: call normalize_delta_status(status) at the existing assignment, keeping the key absent when Status is empty (unchanged). Adds a unit test reusing the baseline_resolved_4c fixture's bracket/ emphasis matrix (bare NEW, [NEW], **[NEW]**, `[NEW]`, UPDATED, UNCHANGED) against parse_threats_findings directly. No golden fixture moves: the golden-generating fixture (agentic_app) resolves to data tier 1, so parse_threats_findings's tier-3 branch is never exercised by test_existing_templates_unchanged. Verified by call-site trace plus an empirical full diff of all 5 template goldens. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Replaces compute_risk_funnel's count-based, mismatched-source funnel
(risk-scores.md counts at tier 1, compensating-controls.md counts at
tiers 2-3) with the volume-based, one-row-set computation specified in
data-model.md §4 (D-2) and contracts/extraction-data-contract.md.
- funnel_tiers is always 4 objects (no null entries), each carrying
ghost, volume (Decimal, 1dp ROUND_HALF_UP) and an integer
severity_mix, split across _funnel_4tier_mode, _funnel_3tier_mode and
_funnel_threats_only_mode per data source tier.
- 4-tier mode: V2 = sum(inherent), V3 = sum(residual if found else
inherent), V4 = sum(residual), over compensating-controls.md rows
with an inherent score. Volumes are unavailable (null, one warning)
when no row carries one, or the quantized V2 is 0.
- 3-tier mode: V2 = sum(composite scores); JSON tier 2 ("Unmitigated
Risk") mirrors V2 exactly, so its reduction is 0.0; JSON tier 3 is
always ghost.
- Threats-only mode: JSON tiers 1-3 are all ghost, plain STEP cascade.
- Widths (FR-K11.5/PD-4): STEP=10, FLOOR=30, Decimal clamp bounds
(never a bare int), round-half-up to int.
- Reductions and risk_reduction (Tier 2->4) are derived from the
emitted, quantized volumes (PD-5's quantize-then-derive), with a
zero-denominator giving 0.0 when volumes are available.
missing_enrichments and control_coverage_pct are untouched legacy
fields, not part of data-model §4 — moved to the call site directly
rather than threaded through compute_risk_funnel's new return shape.
Not in this stage (per the task's explicit scope): S-9's baseball-card
risk_reduction/inherent_score/residual_score wiring, the Section 1
comparand and row-count warnings, and the funnel source strings'
one-row-set wording (N13) — a plain "compensating-controls.md" /
"risk-scores.md" placeholder is used for now.
Verified against all 6 fidelity_373 funnel fixtures (README's
hand-computed arithmetic: STEP-bound, FLOOR-bound, 3-tier, threats-only,
volumes-unavailable, inherent-less join) and independently against
contracts/extraction-data-contract.md's own worked example on the
golden fixture (agentic_app: V2/V3/V4 = 209.8/187.8/167.6, widths
100/90/80/70, reductions 10.5/10.8, risk_reduction 20.1) — exact match
in both cases.
Moves the risk-funnel golden (funnel_tiers shape, reduction_percentages,
risk_reduction, inherent_score, residual_score) — authorized by SC-8 for
K11. The other 4 goldens' only diff is the pre-existing T025 posture
fields; no unauthorized drift.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…021) S-9: the baseball card's risk_reduction/inherent_score/residual_score now take the row-derived values (data-model §4.4) instead of compensating-controls.md Section 1's stated figures. compute_risk_funnel is computed once, shared between the baseball-card and risk-funnel branches of main(), so both surfaces agree by construction and any volumes-unavailable warning prints exactly once per run. Verified null parity in 3-tier, threats-only and volumes-unavailable modes, and value parity against every 4-tier funnel fixture. Warnings (data-model §4.5), both scoped to 4-tier mode in _funnel_4tier_mode: the Section 1 comparand (inherent score, residual score, risk reduction — each warns independently, per field, when it differs from the row-derived value by more than 0.1) and the controls/risk-scores row-count mismatch. Verified against the controls_warnings_kitchen_sink fixture: the stated inherent total and reduction warn, the stated residual (which matches by design) does not, and the 10-vs-11 row-count mismatch warns with the exact contract stem. N13 (align the funnel source strings with the one-row-set wording): verified, no change needed. Every funnel_tiers[].source in the current code already matches contracts/extraction-data-contract.md's worked examples exactly (4-tier: uniformly "compensating-controls.md" since stage 2's rewrite; 3-tier and threats-only: unchanged, already correct). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
_merge_delta_status now delegates to the shared tachi_parsers helpers (delta_status_by_id + apply_delta_status) instead of re-parsing Section 7 with its own raw, non-normalized status read. Keeps its exact (findings, threats_md) signature so the existing test_extractor_contract_fixes.py call site is unaffected — only the resolved status values change, from raw cell text to the normalized form ([NEW], **[NEW]**, `[NEW]` -> NEW) that compute_delta_counts already counts (FR-K12.2, data-model.md §6). main()'s delta-counts block now also calls warn_delta_scope with the selected tier's own finding-ID set, emitting PD-16's scoped, aggregated warnings (unknown statuses, empty map, ID-set mismatch, baseline run with no Status column). The call is unconditional; warn_delta_scope no-ops on its own unless has_baseline. Verified against tests/scripts/fixtures/fidelity_373/baseline_resolved_4c, baseline_resolved_4b_legacy, baseline_status_id_mismatch and non_baseline_no_status_column: delta_counts and the emitted warning text match the fixtures' hand-computed expected values exactly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… tier (T017)
Implements the data-model.md §7 / FR-K13.1 precedence chain so a
recommendation shown in the PDF is never blank, per data tier:
- Tier 1: the finding's `recommendation` already carries the
compensating-controls.md Section 4 analyzer join
(parse_compensating_controls_md). _apply_recommendation_fallback now
falls back to REC_FALLBACK_PREFIX + the threats.md Section 7 Mitigation
column when that's non-empty, else REC_PLACEHOLDER. Mutates the finding
in place so the card (findings-detail.typ:95), the remediation roadmap
and the attack-path remediation (_get_finding_mitigation) all read the
one resolved field. _section4_has_content gates a new drift warning,
checked against the pre-fallback state, when Section 4 has entries but
none of them joined.
- Tier 2: no finding-level field is added (the card keeps its missing-key
default, unchanged). build_remediation_actions' tier-2 branch applies
REC_PLACEHOLDER to the roadmap text only when the threat text itself is
empty. The attack path is unchanged (_get_finding_mitigation already
falls through to _build_remediation's generic step).
- Tier 3: `mitigation` is the only field parse_threats_findings emits;
main() now resolves it to REC_PLACEHOLDER in place when empty, so the
card, the roadmap and the attack path inherit the same text.
REC_FALLBACK_PREFIX ("Threat-model mitigation: ") and REC_PLACEHOLDER
("No recommendation available") are the exact strings from
extraction-data-contract.md.
Verified against tests/scripts/fixtures/fidelity_373/recommendations_partial_join
and recommendations_drifted: per-finding recommendation text and the
drift-warning firing/non-firing both match the fixtures' hand-computed
expected values.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ger (T017) Sweeps the 13 "remaining" Section-4b-meaning-Resolved-Findings references (6 files) to 4c, per FR-K12.5 / spec.md's 16-site/7-file count (the 7th file, scripts/tachi_parsers.py, was already brought into compliance by FR-K12.4/T016 in W1). Leaves every reference that means Findings by Agentic Pattern (Section 4b, still current, untouched) exactly as-is. Full per-occurrence disposition, plus the repo-wide grep for other Resolved-Findings-meaning "4b" sites left out of scope (the non-distributed legacy agents/ tree, the frozen init-baseline-tree fixture, historical specs/PRD artifacts, AOD step numbers, and test infrastructure), is in specs/373-adopter-install-output-fidelity/sweep-4b-ledger.md. Prose/comment-only; no parsing logic changes. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…tribution_refs (T031) _warn_unmatched_attribution_refs's docstring already documented that it never raises and never changes data (true since 3d67ca7, which removed its catalog I/O). This adds the one fact it didn't cover: its sole caller, build_per_framework_aggregates, skips calling it entirely for a framework whose in-scope record count is 0 (the items = [] branch) -- with zero in-scope records there is nothing a stray ref could have been silently dropped against, so the guard would be vacuous, not unsafe to skip. Docstring-only; no behavior change. "Never raises" is unchanged and untouched -- this documents a separate, orthogonal fact about when the caller chooses not to invoke the guard at all. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Emit #let risk-posture-level and #let risk-posture-label right after the severity-count block in report-data.typ, computed by tachi_parsers' compute_risk_posture() on the same post-clamp/inherent/qualitative severity dict that already feeds critical-count/high-count/medium-count/low-count (data-model.md §5, FR-K13.4). This is the report-data.typ half of D-3's single posture rubric; B2a emits the matching metadata.risk_posture_level/ _label into the infographic JSON so the sibling-parity test can compare both surfaces for the same run directory. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ata guard (T025) - cover.typ: cover-page takes risk-posture-level/risk-posture-label from report-data.typ instead of deriving them locally. The removed risk-posture() function's rubric moves to compute_risk_posture() in tachi_parsers.py (previous commit); only the color-from-level mapping stays here, and the label renders verbatim (D-3). - main.typ: a stale-data guard between the _report-data-dict binding and the cover-page call panics with the contract's exact message when a report-data.typ predates the posture variables, rather than silently degrading or re-deriving a second rubric. Wires the two new fields into the cover-page call. - typst-template-contract.md: documents risk-posture-level/-label as REQUIRED, no default — the only variables in the contract without a main.typ fallback. - report-assembler.md: one-line addition to the existing deprecated-path note — a hand-built report-data.typ without the posture variables no longer compiles. Verified locally on Typst 0.14.2 against the mmdc-free posture fixture: a freshly generated report-data.typ compiles and the cover shows "HIGH RISK" (matching the fixture's expected level/label); the same file with the two #let lines stripped panics with the contract's exact message. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
B2b's part of T022 (K11 text): adds the R-P5 clause to the Risk Reduction Funnel page's caption in templates/tachi/security-report/main.typ, clarifying two things a reader could otherwise misread from the image alone: - Tier 3 credits only fully effective (found) controls, while Tier 4 additionally credits partially effective ones (data-model.md §3's "Tier-3 per-row score: residual if status_class == found else inherent" vs. Tier 4's unconditional residual sum) — so the two tiers can diverge even for the same finding set; - bar widths narrow by at least one step per stage for readability (data-model.md §4.1's STEP/FLOOR cascade), while the percentages printed beside them remain exact. The edit is a single new paragraph appended to the existing description block, well clear of stage 3's stale-data guard earlier in the same file. Lane C2 owns the rest of T022 (infographic-risk-funnel.md and template-specific-formats.md) separately. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Add compute_allow_list(template, findings, scope, template_data, payload)
to extract-infographic-data.py (AR-3: not a tachi_parsers.py shared
helper, single consumer). Emits the top-level allow_list key -
{finding_ids, component_names}, sorted/unique/empty-dropped - in every
template's JSON, executive-architecture included through its early-exit
builder (data-model.md §8, PD-17).
finding_ids per template: the full tier finding-ID set for baseball-card
and system-architecture, per_layer_summaries[].top_findings[].id for
maestro-stack, none for maestro-heatmap/risk-funnel, and
callouts[].finding_id for executive-architecture. component_names is one
common set: scope components, trust-zone members and zone names,
data-flow endpoints, and the tier findings' component values.
Per LOW-4, the allow_list lines are kept apart from T025's K13-posture
lines in both _build_executive_architecture_payload and
build_json_output (separate statements, non-adjacent), so a TW-6 revert
of this commit applies cleanly on top of the K13-posture commit.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Add A11 to test_gemini_request_contract.py: the /tachi.infographic command's explicit-path data-source detection condition must accept the short-form "Residual" alias (FR-K9.3), not only "Residual Score". Kept in its own block (banner-separated from where A9/A10 will land in T029) per T039 §8.3, so a later K15 carve's revert never touches A11. Written test-first from contracts/gemini-request-and-scaffold.md's Static contract test table and spec.md FR-K9.3 / US-3a #10. Scoped to the detection CONDITION line only, not the "UNABLE TO DETECT DATA SOURCE TYPE" fallback message's prose: spec.md's "Header drift" edge case scopes FR-K9.3 to "the infographic command's tier detection" (the condition deciding whether a short-form file is rejected at all), and that fallback message only fires when no source type matches any indicator -- a valid short-form file never reaches it once the condition accepts both aliases. An earlier draft also asserted the fallback message's wording and stayed red against the merged probe's T018 (373-w2-C2, e2440d5); trimmed after re-reading FR-K9.3 rather than left pinned to a requirement the contract does not actually make. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Extractor-level regression tests for US-3a (K9/K10/K12/K13.1), written
test-first from contracts/extraction-data-contract.md, data-model.md
§3/§6/§7 and the fidelity_373 fixture README's hand-computed expected
values. These exercise extract-infographic-data.py's and
extract-report-data.py's own wiring (extract_severity,
parse_maestro_layer_distribution / parse_maestro_data, the CLI's
delta_counts and recommendation-resolution call sites) -- distinct from
test_tachi_parsers.py's T016 pins, which already cover the shared
tachi_parsers.py functions directly.
test_extract_infographic_data.py (+6): K9 (short-form bands wire through
extract_severity), K10 (a "###" heading and its "####" twin produce
byte-identical parse_maestro_layer_distribution output), and K12 (4c
exact counts with bracketed statuses, 4b legacy parity, non-baseline
no-warning, and the F1/NM-1 ID-mismatch case -- absolute tallies plus
exactly one warning).
test_extract_report_data.py (+13): the K9/K10/K12 mirror of the above
(parse_compensating_controls_md's import binding, parse_maestro_data
equivalence, delta_counts via report-data.typ), plus K13.1's per-tier
recommendation resolution: tier-1 precedence (analyzer -> prefixed
Section 7 mitigation -> placeholder) and drift warning, the roadmap's
tier-2 placeholder guard, the tier-2 attack-path anchor that must NOT
show the placeholder (data-model.md §7's one documented divergence
between the roadmap and the attack path), and a tier-3 end-to-end check
that the empty-mitigation placeholder is resolved once and read
identically by the card, the roadmap and the attack path. Adds fixture
recommendations_tier3_empty_mitigation/ (documented in the fixture
README) since T003's list only covered tier-1 K13.1 scenarios.
test_extractor_contract_fixes.py (+2): K13.1 consumer-parity anchors --
"on tiers 1 and 3 the three consumers read one field and so always
agree" (data-model.md §7), pinned independently of T017's exact
resolution logic so a future change can never give one consumer its own
separate path.
Verified against a scratch probe merging the three W2 implementation
lanes (373-w2-B2a, 373-w2-B2b, 373-w2-C2): all tests in this commit now
pass there. Two tests were corrected during that verification, not bent
to fit the implementation -- re-reading the contract showed the original
assertions were the ones in error:
- the tier-2/tier-3 roadmap placeholder guards: the tier-3 case assumed
build_remediation_actions must apply the placeholder itself; B2b
instead resolves it once upstream in main() per data-model.md §7's
literal wording ("the finding's mitigation, so the placeholder when
empty" describes reading one resolved field, not a second guard).
Replaced the tier-3 unit-level call with the end-to-end check above.
- a fixture-README transcription of two Section 7 Mitigation cells added
a trailing period the source markdown does not contain; corrected both
the assertions and the README prose to the verbatim source text.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…9, W2 wiring) Cherry-picked from the test lane (9a690b0) without its own commit: tests/scripts/test_extraction_sibling_parity.py compares posture, severity counts, delta counts (including bracketed and ID-mismatch baselines) and MAESTRO output between the infographic JSON and report-data.typ on the same fixtures. Wires the module into the extraction-fidelity job's pytest invocation and adds it to the shared paths: anchor. Its fixture trees (report_data/** and fidelity_373/**) are already covered since A-2 (C-1), so no further paths: growth is needed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds the two tests tasks.md T031 specifies, retargeted to today's code (commit 3d67ca7 moved the guard out of classify_framework_items into its own _warn_unmatched_attribution_refs, called from build_per_framework_aggregates). OQ-5 is closed: Case 2 uses the real schemas/taxonomy/*.yaml catalogs. - Case 1: a stale-form attribution id (an old year suffix against a newer catalog) warns on stderr; classify_framework_items's output and the caller's findings list are both unaffected by the guard call (LOW-7 purity). - Case 2: through build_per_framework_aggregates with the real catalogs, a finding citing mitre-attack T1070.001 (out_of_scope: true at schemas/taxonomy/mitre-attack.yaml:982-988) is not reported as unmatched -- the guard's catalog_ids set is the full catalog, not the in-scope filter classification uses. Two tests only, per the recipe's TW-4 cap. Both pass against today's code. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
… (T025, W2 wiring) Cherry-picked from B2b (1c8b6d5) without its own commit: tests/scripts/test_report_posture_contract.py proves a stale report-data.typ (regenerated from an mmdc-free fidelity_373 fixture, never from examples/) panics with the regenerate message in main.typ's guard, and that the positive control compiles with no panic text. Adds a new, stand-alone report-posture job: ubuntu-latest, typst-community/setup-typst@v5 (unpinned, matching tachi-mmdc-preflight.yml's precedent), TACHI_REQUIRE_TYPST=1 so the gate fails rather than skips when Typst is missing, running only this module. Adds templates/tachi/security-report/** and the module to the shared paths: anchor. This commit holds only this job and this module (PD-8): if K13-posture is ever carved, dropping this job and its two paths: entries is the whole revert. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds extractor-level (CLI end-to-end) coverage for K11's risk funnel (compute_risk_funnel, data-model.md §4) in tests/scripts/test_extract_infographic_data.py, kept in its own contiguous block after the K9/K10/K12 section so a K11-only revert applies cleanly: - STEP-bound (funnel_step_bound) and FLOOR-bound (funnel_strong_reduction) widths, volumes and reductions. - The 3-tier, threats-only and volumes-unavailable shapes (ghost flags, widths, null semantics), matching contracts/extraction-data-contract.md. - Control-status variants (Missing/empty/unrecognized/Partially Found/ None found) and the clamp, wired through the funnel's row set via controls_warnings_kitchen_sink -- not just at classify_control_status's own parser-level unit (test_tachi_parsers.py/T020). - The Section 1 comparand (two per-field warnings, never aggregated, the matching field silent) and the row-count-mismatch warning -- both T021 follow-on work, matched against the contract's stem rather than a guessed <field> spelling. - The Inherent-less join-path fixture: volumes come from the risk-scores composites joined by ID. - S-9 baseball-card/funnel parity on risk_reduction, inherent_score and residual_score, including the null case (volumes unavailable). - Reductions are never negative across the funnel fixture set. Values are hand-computed in tests/scripts/fixtures/fidelity_373/README.md. Expected red until Lane B2a's S-9 totals/warnings work lands on top of its already-merged K11 funnel computation (T021, a6b889f). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds the funnel_join_inherent_less/ case (K11 carve unit; architect finding F2) to test_extraction_sibling_parity.py: on a Coverage Matrix with no Inherent Score/Inherent column at all, both rows' inherent scores come only from the sibling risk-scores.md composites joined by ID. Asserts the clamp-processed residual bands (T-1 7.5 -> High, T-2 6.6 -> Medium), the resulting severity counts, and the posture agree between the infographic JSON and report-data.typ -- both extractors' tier-1 call sites read risk-scores.md for this join (T020, W1-landed on both sides), so this proves one shared pipeline, not two independent implementations of it. The severity-count test passes today. The posture test is expected red until T025 lands on both Lane B2a and Lane B2b, matching the existing posture_mmdc_free precedent in this file -- not weakened, skipped, or xfail-marked. Updates the module docstring's fixture list and removes the stale "not covered here" note now that this case is covered. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds A9 (layout-label + allow-list prompt hardening) and A10 (executive-architecture lock amendment) to tests/scripts/test_gemini_request_contract.py, in their own block between A8 and A11 -- separated from A11 by that class's own banner, matching the layout the file already documented, so a K15 carve (TW-6, tasks.md T030) reverting T027-T029 removes exactly this block. A9: the layout-label instruction and the allow-list rule (or, for executive-architecture, its CALLOUTS/LAYER STACK/FLOW EDGES/CLUSTERS region variant) both sit between IMPORTANT: and STYLING DIRECTIVES in each of the five scaffolded preambles and the reference prompt; the agent text instructs writing the ALLOWED IDS AND NAMES line from allow_list for those six surfaces; an N10 negative confirms the agent's executive-architecture section does NOT re-instruct that line. All literal K15 text is quoted verbatim from the contract's "K15 instruction text" section and independently re-verified byte-for-byte against the live template/reference/agent files before being hard-coded (not trusted from contract prose alone). A10: the lock-marker lines are byte-identical, flow_edges/clusters are still present, the dated "Amended by F-373 K15" note exists in the reference, and -- condition 9 -- the locked block with the amendment paragraph (the one immediately after IMPORTANT:) removed hashes to the pinned pre-K15 SHA-256. The pin (12db7d757046fb0f38af6a9940b9403319ef11ae84144010803959d1f914eb38) was independently recomputed here from `git show 63438d7:.claude/skills/tachi-infographics/references/executive-architecture.md` per the contract's exact extraction/normalization rule (L13) and matches Lane C2's own computation of the same value. Verified: all 13 new tests are red against this branch's own (pre-K15) files -- equivalent to 2b59e05's text, since no K15 commit has landed here -- and green when run against Lane C2's tip (3151c05) via a temporary copy into that worktree (reverted after, worktree left clean). A1-A8 continue to pass unmodified; A11 remains red here for the same pre-existing, unrelated reason stage 1 already recorded (C2's T018 not yet merged into this branch). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Regenerate the K11-authorized golden fields (SC-8) from the tier-1
agentic_app fixture, per quickstart.md §5. Fixture command:
python3 scripts/extract-infographic-data.py \
--target-dir tests/scripts/fixtures/exec_arch/agentic_app \
--template <t> --output <out>.json
Leaves patched onto the committed goldens (K11 class only: funnel_tiers[]
ghost/severity_mix/volume/width/source, inherent_score, residual_score,
risk_reduction, reduction_percentages — the row-derived risk computation
from T020/T021):
risk-funnel.json (24 leaves):
template_data.funnel_tiers[0].{ghost,severity_mix,volume,width}
template_data.funnel_tiers[1].{ghost,severity_mix,source,volume,width}
template_data.funnel_tiers[2].{ghost,severity_mix,volume,width}
template_data.funnel_tiers[3].{ghost,severity_mix,source,volume,width}
template_data.inherent_score
template_data.reduction_percentages[0..2].percentage
template_data.residual_score
template_data.risk_reduction
baseball-card.json (3 leaves, S-9):
template_data.inherent_score
template_data.residual_score
template_data.risk_reduction
maestro-heatmap.json, maestro-stack.json, system-architecture.json:
unchanged (no K11 fields).
Verified leaf-by-leaf (type-strict) against a scratch regeneration:
no leaf outside this class differs at this stage.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Regenerate the K13-authorized posture fields (SC-8) onto the K11-patched
goldens, from the tier-1 agentic_app fixture, per quickstart.md §5.
Leaves patched (K13-posture class only: metadata.risk_posture_level and
metadata.risk_posture_label, computed once by compute_risk_posture on
post-clamp severity counts, T024/T025), 2 leaves in every file:
risk-funnel.json, baseball-card.json, maestro-heatmap.json,
maestro-stack.json, system-architecture.json:
metadata.risk_posture_label ("HIGH RISK")
metadata.risk_posture_level ("high")
Verified leaf-by-leaf (type-strict) against a scratch regeneration:
exactly these 2 leaves per file changed from the K11 stage; no leaf
outside this class differs at this stage.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Regenerate the K15-authorized prompt_scaffold and allow_list fields
(SC-8) onto the K13-patched goldens, from the tier-1 agentic_app
fixture, per quickstart.md §5.
Leaves patched (K15 class only: the layout-label sentence and the
ALLOWED IDS AND NAMES rule added to each scaffolded preamble between
IMPORTANT and STYLING DIRECTIVES per T028, and the new top-level
allow_list emitted by compute_allow_list per T027), 2 leaves in every
file:
risk-funnel.json, baseball-card.json, maestro-heatmap.json,
maestro-stack.json, system-architecture.json:
allow_list (new top-level key: component_names,
finding_ids)
prompt_scaffold.preamble (K15 sentence inserted)
prompt_scaffold.postamble is byte-identical to the prior committed
value in every file, confirmed before patching (SC-8 requirement).
After this commit, every golden in tests/scripts/fixtures/golden/ is
byte-identical (cmp) to a direct regeneration of
scripts/extract-infographic-data.py against
tests/scripts/fixtures/exec_arch/agentic_app in a scratch clone of
this commit's parent tree — verified leaf-by-leaf (type-strict) at
each of the three stages, with no leaf outside the K11/K13/K15 classes
differing anywhere (no delta_status, severity-count, or top_findings
drift).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…25-T029, T031, T032) W2 integrated locally from five lane worktrees plus two lock-step wiring commits (sibling parity in extraction-fidelity; the K13-posture report-posture job). Goldens regenerated as one commit per K-item (K11 5a1810e, K13 3ffc3db, K15 982c074; 47 authorized leaves). Fast set 254 passed, 0 failed. T016's tier-3 normalization gap was found by B2b and fixed in W2 (9019528). SEC-K3-01/02 from T011's review fixed in 45bb8d6 pending the architect's P0 ratification. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…on 1 handoff P0 (architect): parser semantics conform; goldens approved (47 leaves); 9019528 ratified as a scoped AR-3 exception; the 32-hop ceiling ratified; installer :145, data-model and extraction contract amended at P0. Three required changes are due in W3 before T036: RC-1 (MEDIUM, K3 resolve() existence guard for the multi-ancestor partial write), RC-2 and RC-3 (LOW). SEC-K3-03 goes to a follow-up issue. N4 re-run: totals identical to W0; maestro-reference +2 PDF pages is K13.1 M5, intended; baselines untouched (NFR-8). TW-7 (team-lead): not fired (1.72 d vs 2.0; max lane 47%), so T040 is marked not triggered. Wave-3 gated set 482 passed, 0 failed; CI green on all workflows at 4bea6c5. NEXT-SESSION.md is the Session 2 handoff. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… with zero writes (P0 RC-1) Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…2) (P0 RC-3, K9)
The stop rule was `level <= matched_level`, which loosens a level-1
match: a bare-substring fallback ("Risk Summary", "Severity
Distribution", "Coverage Distribution") that happens to match a
document TITLE line (itself a level-1 heading) would then scan past
the next "##" section and could adopt a later section's table,
breaking FR-K9.2's "existing callers MUST be unaffected". The
docstring's claim that "same or higher level" reduces exactly to the
old hardcoded rule for a level-1-or-2 match was false for level 1.
Change the stop rule to `level <= max(matched_level, 2)`, so a
level-1 match keeps the old "#"/"##" stop like a level-2 match
always has, and only a level-3+ match gets the tighter, level-aware
rule. Correct the docstring and the nearby stale test-file comment
making the same false generalization. Add a regression test pinning
the level-1 title case ("# Threat Model Risk Summary" matched by
"Risk Summary", followed by "## 1. Components") to [] (W0 behavior).
No tracked example hits a level-1 line, so this changes no oracle or
golden output.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…unt warning only on a readable table (P0 RC-2, K11)
_funnel_4tier_mode re-parsed risk-scores.md a second time just to
count rows, after extract_severity's tier-1 branch had already
parsed it once to build the K11 inherent-score join. When the Scored
Threat Table couldn't be found, this doubled the "could not find
Scored Threat Table in risk-scores.md" warning and printed a false
"controls rows (N) differ from risk-scores rows (0)" comparison —
the risk-scores row count was unknown, not zero. Seen on
maestro-reference and consumer-agent-app/sample-report in every
baseball-card and risk-funnel run.
Parse risk-scores.md once, in extract_severity's tier-1 branch, and
carry the row count forward on cc_data ("risk_scores_row_count";
None when risk-scores.md is absent or its table is unreadable).
_funnel_4tier_mode now reads that field instead of re-parsing, and
only compares row counts when it is not None (i.e. the parse yielded
at least one row) — so it drops its now-unused rs_content parameter.
compute_risk_funnel's call site and docstring (which already
documented rs_content as "only consulted when tier == 2") follow.
Added a regression test reusing the kitchen-sink fixture's
threats.md/compensating-controls.md with a deliberately mis-headed
risk-scores.md ("## Section 2: Scored Threat Table"): exactly one
"could not find" warning, no false row-count comparison. The
existing 10-vs-11 genuine-mismatch warning test still fires. No
golden or JSON change.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ve() (P0 RC-1, SEC-K3-01 residual) Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Draft PR for feature 373 (bundle lead issue #373, defects K1–K3 and K9–K15). Opened automatically at plan stage.
docs/product/02_PRD/373-adopter-install-output-fidelity-2026-09-27.md(Approved)specs/373-adopter-install-output-fidelity/feasibility-check.mdfix(373)), per PRD ruling C-8Closes #373
🤖 Generated with Claude Code