cockpit-agent/README.md

135 lines
4.7 KiB
Markdown
Raw Permalink Normal View History

2026-08-24 16:49:44 +08:00
# cockpit-agent
`cockpit-agent` implements a single-pass automotive cockpit task pipeline:
```text
instruction
-> SemanticAction
-> function semantic target
-> function grounding
-> ROI crop
-> UIState
-> ActionPlan
-> action grounding (when needed)
-> MockAction
```
It owns intent parsing, current UI understanding, deterministic planning, and
mock action generation. The existing `cockpit-ui-grounding` package owns:
```text
image + target -> bbox -> pixel (u, v)
```
The current stage is deliberately limited to **NO ROBOT**, **NO FEEDBACK**, and
**NO CLOSED LOOP**. It does not connect to ROS, a mechanical arm, or any real
touch API, and it does not verify, retry, or replan an action.
## Setup
All model paths in `configs/default.toml` are local. No network model ID is
used.
```bash
conda activate qwen3vl
pip install -e /data/lgv/projects/cockpit-ui-grounding
pip install -e /data/lgv/projects/cockpit-agent
export CUDA_VISIBLE_DEVICES=0
```
The default configuration uses one shared Qwen3.5-2B instance for intent, UI
understanding, and grounding. `ModelRegistry` caches models by backend and
resolved local path, so identical settings do not load duplicate weights.
Function ROI padding is configured independently with
`perception.roi_padding_ratio`; the default is `0.15`.
## Run A Task
```bash
python scripts/run_task.py \
--image /data/lgv/datasets/cockpit-ui/raw/debug/ca_car.jpg \
--instruction "把主驾温度调到23度" \
--output-dir /data/lgv/runs/cockpit-agent/demo_driver_temp_23
```
Focused CLIs are also available:
```bash
python scripts/run_intent.py --instruction "打开内循环"
python scripts/run_ui_understanding.py \
--image /data/lgv/datasets/cockpit-ui/raw/debug/ca_car.jpg \
--instruction "后备箱最大开度设置成80%" \
--output-dir /data/lgv/runs/cockpit-agent/ui_understanding
```
A successful task output directory contains:
- `intent.json`
- `function_grounding.json`
- `roi.jpg`
- `ui_state.json`
- `plan.json`
- `action_grounding.json` when a non-noop action is proposed
- `action.json`
- `result.jpg`
- `run.json`
The trace uses explicit stage names: `semantic_action`, `function_target`,
`function_grounding`, `roi`, `ui_state`, `action_plan`, `action_target`,
`action_grounding`, and `mock_action`. Reusing an output directory replaces
the pipeline-owned artifacts so a failed run cannot retain stale results from
an earlier run.
If a stage raises an error, the pipeline writes `failure.json` and a partial
`run.json` with `success`, `failed_stage`, `error_type`, `error_message`, and
`raw_model_output`. It then re-raises the error so callers cannot mistake a
failed run for a successful one.
`action.json` is always a proposal. A tap contains a pixel and repeat count, a
drag contains start/end pixels, and a no-op contains its reason. Nothing is
executed outside the process.
## Model Output Handling
Intent and ROI UI model replies are parsed as JSON and validated against
strict dataclass schemas before planning. A small deterministic normalization layer
handles known representation-only differences such as `"100%"` versus `100`
and Chinese unlock option labels versus `all_doors`; it does not invent missing
visual state. UI observation fields are nullable. Slider observations separate
endpoint values from `current_value` and expose nullable `track_visible` and
`knob_visible` fields; an unobservable current value remains `null` rather than
being guessed. The first implementation proposes a drag only when the model
explicitly observes `orientation=horizontal`; unknown and vertical sliders fail
instead of silently using horizontal geometry. Device state (`current_state`) is separate from visual choice
state (`selection_state`). Common choice-state wording such as
`active`/`selected` and `inactive`/`unselected` is normalized without changing
the model's predicted control type.
`SemanticTargetBuilder` creates a complete-control/group query before UI
understanding and a specific action-control query after planning. The function
bbox is padded by the configured ratio and cropped without resizing. ROI action
boxes are converted back to source-camera pixels. Slider drags use the visually
grounded knob center as the start and the grounded track plus
`normalized_target` as the end. The adapter does not repair predictions with
image-, vehicle-, color-, position-, or fixed-ROI rules.
The benchmark integrity policy is documented in
[`docs/generalization_rules.md`](docs/generalization_rules.md). In particular,
unit tests validate deterministic program logic, while model benchmarks may
fail and must report the model's real output.
## Tests
Planner tests are deterministic and do not load a model:
```bash
pytest -v
```
The tests also run with the standard library runner:
```bash
python -m unittest discover -s tests -v
```