54 lines
2.5 KiB
Markdown
54 lines
2.5 KiB
Markdown
# Generalization and Benchmark Integrity
|
|
|
|
The purpose of the automotive UI benchmark is to measure model capability.
|
|
A model prediction is allowed to fail and must remain visible in the artifacts.
|
|
|
|
## Architecture Boundary
|
|
|
|
Business logic may be deterministic. Examples include calculating a stepper's
|
|
repeat count, deciding that a toggle is already in the requested state, and
|
|
normalizing a slider value from its observed minimum and maximum.
|
|
|
|
Visual layout may not be deterministic. Planner and adapters must not infer a
|
|
target from a vehicle brand, image filename, benchmark ID, color, neighboring
|
|
control, screen region, fixed pixel, fixed bounding box, crop, or ROI. Planner
|
|
operates on semantic functions and never on visual positions.
|
|
|
|
`SemanticTargetBuilder` may convert structured semantic fields into natural
|
|
language such as `driver temperature decrease control`. It must describe only
|
|
the target's semantic identity. The grounding model is solely responsible for
|
|
finding that target in the current image.
|
|
|
|
Every perception ROI is derived from the function grounding bbox. Padding may
|
|
be a configurable ratio of that bbox; a fixed screen crop is prohibited. UI
|
|
Understanding receives only this dynamically generated ROI. Grounding inside
|
|
the ROI is translated back to source-camera coordinates before a MockAction is
|
|
created.
|
|
|
|
Unobservable UI fields remain null. In particular, slider motion starts from a
|
|
visually grounded knob, never from a guessed `current_value`.
|
|
|
|
## Dataset Policy
|
|
|
|
- Development data may be inspected while improving models, training data, or
|
|
general prompts.
|
|
- Validation data is used for model selection.
|
|
- Test data is frozen and must never be used to add matching rules.
|
|
- No image-specific, vehicle-specific, or benchmark-case-specific branch is
|
|
permitted in production code.
|
|
|
|
## Test Policy
|
|
|
|
Unit tests validate deterministic schemas, parsing, planning, semantic query
|
|
construction, coordinate conversion, and failure recording. They do not assert
|
|
that a vision model must return a hand-observed box for a fixed image.
|
|
|
|
Model benchmarks report semantic correctness separately for Intent, UI
|
|
Understanding, Grounding, and the proposed MockAction. A schema-valid but
|
|
semantically wrong bounding box is a benchmark failure, not a reason to add a
|
|
postprocessing rule.
|
|
|
|
Model capability problems should be addressed with representative automotive
|
|
data, model training or fine-tuning, and evaluated model/prompt improvements.
|
|
They must not be hidden with layout-specific code.
|