# cockpit-agent `cockpit-agent` implements a single-pass automotive cockpit task pipeline: ```text instruction -> SemanticAction -> function semantic target -> function grounding -> ROI crop -> UIState -> ActionPlan -> action grounding (when needed) -> MockAction ``` It owns intent parsing, current UI understanding, deterministic planning, and mock action generation. The existing `cockpit-ui-grounding` package owns: ```text image + target -> bbox -> pixel (u, v) ``` The current stage is deliberately limited to **NO ROBOT**, **NO FEEDBACK**, and **NO CLOSED LOOP**. It does not connect to ROS, a mechanical arm, or any real touch API, and it does not verify, retry, or replan an action. ## Setup All model paths in `configs/default.toml` are local. No network model ID is used. ```bash conda activate qwen3vl pip install -e /data/lgv/projects/cockpit-ui-grounding pip install -e /data/lgv/projects/cockpit-agent export CUDA_VISIBLE_DEVICES=0 ``` The default configuration uses one shared Qwen3.5-2B instance for intent, UI understanding, and grounding. `ModelRegistry` caches models by backend and resolved local path, so identical settings do not load duplicate weights. Function ROI padding is configured independently with `perception.roi_padding_ratio`; the default is `0.15`. ## Run A Task ```bash python scripts/run_task.py \ --image /data/lgv/datasets/cockpit-ui/raw/debug/ca_car.jpg \ --instruction "把主驾温度调到23度" \ --output-dir /data/lgv/runs/cockpit-agent/demo_driver_temp_23 ``` Focused CLIs are also available: ```bash python scripts/run_intent.py --instruction "打开内循环" python scripts/run_ui_understanding.py \ --image /data/lgv/datasets/cockpit-ui/raw/debug/ca_car.jpg \ --instruction "后备箱最大开度设置成80%" \ --output-dir /data/lgv/runs/cockpit-agent/ui_understanding ``` A successful task output directory contains: - `intent.json` - `function_grounding.json` - `roi.jpg` - `ui_state.json` - `plan.json` - `action_grounding.json` when a non-noop action is proposed - `action.json` - `result.jpg` - `run.json` The trace uses explicit stage names: `semantic_action`, `function_target`, `function_grounding`, `roi`, `ui_state`, `action_plan`, `action_target`, `action_grounding`, and `mock_action`. Reusing an output directory replaces the pipeline-owned artifacts so a failed run cannot retain stale results from an earlier run. If a stage raises an error, the pipeline writes `failure.json` and a partial `run.json` with `success`, `failed_stage`, `error_type`, `error_message`, and `raw_model_output`. It then re-raises the error so callers cannot mistake a failed run for a successful one. `action.json` is always a proposal. A tap contains a pixel and repeat count, a drag contains start/end pixels, and a no-op contains its reason. Nothing is executed outside the process. ## Model Output Handling Intent and ROI UI model replies are parsed as JSON and validated against strict dataclass schemas before planning. A small deterministic normalization layer handles known representation-only differences such as `"100%"` versus `100` and Chinese unlock option labels versus `all_doors`; it does not invent missing visual state. UI observation fields are nullable. Slider observations separate endpoint values from `current_value` and expose nullable `track_visible` and `knob_visible` fields; an unobservable current value remains `null` rather than being guessed. The first implementation proposes a drag only when the model explicitly observes `orientation=horizontal`; unknown and vertical sliders fail instead of silently using horizontal geometry. Device state (`current_state`) is separate from visual choice state (`selection_state`). Common choice-state wording such as `active`/`selected` and `inactive`/`unselected` is normalized without changing the model's predicted control type. `SemanticTargetBuilder` creates a complete-control/group query before UI understanding and a specific action-control query after planning. The function bbox is padded by the configured ratio and cropped without resizing. ROI action boxes are converted back to source-camera pixels. Slider drags use the visually grounded knob center as the start and the grounded track plus `normalized_target` as the end. The adapter does not repair predictions with image-, vehicle-, color-, position-, or fixed-ROI rules. The benchmark integrity policy is documented in [`docs/generalization_rules.md`](docs/generalization_rules.md). In particular, unit tests validate deterministic program logic, while model benchmarks may fail and must report the model's real output. ## Tests Planner tests are deterministic and do not load a model: ```bash pytest -v ``` The tests also run with the standard library runner: ```bash python -m unittest discover -s tests -v ```