# Generalization and Benchmark Integrity The purpose of the automotive UI benchmark is to measure model capability. A model prediction is allowed to fail and must remain visible in the artifacts. ## Architecture Boundary Business logic may be deterministic. Examples include calculating a stepper's repeat count, deciding that a toggle is already in the requested state, and normalizing a slider value from its observed minimum and maximum. Visual layout may not be deterministic. Planner and adapters must not infer a target from a vehicle brand, image filename, benchmark ID, color, neighboring control, screen region, fixed pixel, fixed bounding box, crop, or ROI. Planner operates on semantic functions and never on visual positions. `SemanticTargetBuilder` may convert structured semantic fields into natural language such as `driver temperature decrease control`. It must describe only the target's semantic identity. The grounding model is solely responsible for finding that target in the current image. Every perception ROI is derived from the function grounding bbox. Padding may be a configurable ratio of that bbox; a fixed screen crop is prohibited. UI Understanding receives only this dynamically generated ROI. Grounding inside the ROI is translated back to source-camera coordinates before a MockAction is created. Unobservable UI fields remain null. In particular, slider motion starts from a visually grounded knob, never from a guessed `current_value`. ## Dataset Policy - Development data may be inspected while improving models, training data, or general prompts. - Validation data is used for model selection. - Test data is frozen and must never be used to add matching rules. - No image-specific, vehicle-specific, or benchmark-case-specific branch is permitted in production code. ## Test Policy Unit tests validate deterministic schemas, parsing, planning, semantic query construction, coordinate conversion, and failure recording. They do not assert that a vision model must return a hand-observed box for a fixed image. Model benchmarks report semantic correctness separately for Intent, UI Understanding, Grounding, and the proposed MockAction. A schema-valid but semantically wrong bounding box is a benchmark failure, not a reason to add a postprocessing rule. Model capability problems should be addressed with representative automotive data, model training or fine-tuning, and evaluated model/prompt improvements. They must not be hidden with layout-specific code.