◆ NFRGate / Rule Reference

L2 — Structured error logs

medium 📝 logs · both

Logs emitted from an error path use structured fields (not a bare or unformatted string) and include at minimum: error type and originating operation/function.

Python Implementation

L2 (static component): error-path logs use structured fields, not a bare
or unformatted string. rubric_store/definitions/logs.yaml tags L2 `both` —
static analysis can confidently catch the two obvious shapes ("no log call
at all" and "log call with only a bare/f-string message, no structured
fields") but genuinely ambiguous cases are still LLM/Sprint-3 territory.
This rule only emits the high-confidence obvious cases; it does not attempt
the full criterion.

Both fail-branch confidences retuned from 0.85 to 0.778, per
docs/false_positive_rate_study_v1.md: pooled across both branches (they
assert the same underlying claim — "no genuine structured log exists here"
— just reached via different code shapes) and across the Go
implementation's equivalent branches (n=5 total; Java's L2 wasn't sampled).
Pass is untouched (not sampled).

Java Implementation

L2 for Java: catch block's log call uses structured placeholder args
(SLF4J/Log4j style: logger.error("failed: {}", amount, e)), not a single
concatenated/bare string. See ast_helpers.is_structured_log_call.

The "no log call found at all" fail branch's confidence was retuned from
0.85 to 0.08 in docs/product_readiness_report_v1.md's follow-up pass — a
real, hand-labeled sample (n=21, all 21 false positives, Beta(2,2) posterior
mean (0+2)/(21+4)) against real code (FasterXML/jackson-databind), the
first time this branch was ever checked against anything but this
project's own test fixtures. Deliberately NOT pooled with Python/Go's L2
value (0.778, docs/confidence_retuning_v1.md): that sample was real
*service* code (spring-petclinic, fastapi), where "no log in a catch
block" usually is a real gap; this sample is real *library* code, where
the dominant real pattern turned out to be deliberate wrap-and-rethrow via
a helper (or the addSuppressed resource-cleanup idiom) specifically so a
caller-side boundary logs it once instead of every intermediate layer
logging the same failure — a different population, not just a different
language, so pooling them would hide a real distinction rather than
generalize a shared one. Worth a broader look: this may mean the rule
itself should recognize wrap-and-rethrow as an acceptable non-fail pattern
generally, not just a Java-specific confidence adjustment — logged as a
follow-up, not done in this pass (same posture the original false-positive
study took on T1/R1's heuristic redesign).

Go Implementation

L2 for Go: the `if err != nil { ... }` block (constructs.py maps this to
the same "catch_block" kind Python/Java's except/catch produce) contains a
structured (key-value) log call, not just a printf-style message. See
ast_helpers.is_structured_log_call.

Both fail-branch confidences retuned to 0.778 — see structured_log_rule.py
(Python) for the full rationale from docs/false_positive_rate_study_v1.md
(n=4 of the pooled n=5 came from this Go implementation).
← All rules