cockpit-agent/docs/generalization_rules.md

54 lines
2.5 KiB
Markdown
Raw Permalink Normal View History

2026-08-24 16:49:44 +08:00
# Generalization and Benchmark Integrity
The purpose of the automotive UI benchmark is to measure model capability.
A model prediction is allowed to fail and must remain visible in the artifacts.
## Architecture Boundary
Business logic may be deterministic. Examples include calculating a stepper's
repeat count, deciding that a toggle is already in the requested state, and
normalizing a slider value from its observed minimum and maximum.
Visual layout may not be deterministic. Planner and adapters must not infer a
target from a vehicle brand, image filename, benchmark ID, color, neighboring
control, screen region, fixed pixel, fixed bounding box, crop, or ROI. Planner
operates on semantic functions and never on visual positions.
`SemanticTargetBuilder` may convert structured semantic fields into natural
language such as `driver temperature decrease control`. It must describe only
the target's semantic identity. The grounding model is solely responsible for
finding that target in the current image.
Every perception ROI is derived from the function grounding bbox. Padding may
be a configurable ratio of that bbox; a fixed screen crop is prohibited. UI
Understanding receives only this dynamically generated ROI. Grounding inside
the ROI is translated back to source-camera coordinates before a MockAction is
created.
Unobservable UI fields remain null. In particular, slider motion starts from a
visually grounded knob, never from a guessed `current_value`.
## Dataset Policy
- Development data may be inspected while improving models, training data, or
general prompts.
- Validation data is used for model selection.
- Test data is frozen and must never be used to add matching rules.
- No image-specific, vehicle-specific, or benchmark-case-specific branch is
permitted in production code.
## Test Policy
Unit tests validate deterministic schemas, parsing, planning, semantic query
construction, coordinate conversion, and failure recording. They do not assert
that a vision model must return a hand-observed box for a fixed image.
Model benchmarks report semantic correctness separately for Intent, UI
Understanding, Grounding, and the proposed MockAction. A schema-valid but
semantically wrong bounding box is a benchmark failure, not a reason to add a
postprocessing rule.
Model capability problems should be addressed with representative automotive
data, model training or fine-tuning, and evaluated model/prompt improvements.
They must not be hidden with layout-specific code.