135 lines
4.7 KiB
Markdown
135 lines
4.7 KiB
Markdown
|
|
# cockpit-agent
|
||
|
|
|
||
|
|
`cockpit-agent` implements a single-pass automotive cockpit task pipeline:
|
||
|
|
|
||
|
|
```text
|
||
|
|
instruction
|
||
|
|
-> SemanticAction
|
||
|
|
-> function semantic target
|
||
|
|
-> function grounding
|
||
|
|
-> ROI crop
|
||
|
|
-> UIState
|
||
|
|
-> ActionPlan
|
||
|
|
-> action grounding (when needed)
|
||
|
|
-> MockAction
|
||
|
|
```
|
||
|
|
|
||
|
|
It owns intent parsing, current UI understanding, deterministic planning, and
|
||
|
|
mock action generation. The existing `cockpit-ui-grounding` package owns:
|
||
|
|
|
||
|
|
```text
|
||
|
|
image + target -> bbox -> pixel (u, v)
|
||
|
|
```
|
||
|
|
|
||
|
|
The current stage is deliberately limited to **NO ROBOT**, **NO FEEDBACK**, and
|
||
|
|
**NO CLOSED LOOP**. It does not connect to ROS, a mechanical arm, or any real
|
||
|
|
touch API, and it does not verify, retry, or replan an action.
|
||
|
|
|
||
|
|
## Setup
|
||
|
|
|
||
|
|
All model paths in `configs/default.toml` are local. No network model ID is
|
||
|
|
used.
|
||
|
|
|
||
|
|
```bash
|
||
|
|
conda activate qwen3vl
|
||
|
|
pip install -e /data/lgv/projects/cockpit-ui-grounding
|
||
|
|
pip install -e /data/lgv/projects/cockpit-agent
|
||
|
|
export CUDA_VISIBLE_DEVICES=0
|
||
|
|
```
|
||
|
|
|
||
|
|
The default configuration uses one shared Qwen3.5-2B instance for intent, UI
|
||
|
|
understanding, and grounding. `ModelRegistry` caches models by backend and
|
||
|
|
resolved local path, so identical settings do not load duplicate weights.
|
||
|
|
Function ROI padding is configured independently with
|
||
|
|
`perception.roi_padding_ratio`; the default is `0.15`.
|
||
|
|
|
||
|
|
## Run A Task
|
||
|
|
|
||
|
|
```bash
|
||
|
|
python scripts/run_task.py \
|
||
|
|
--image /data/lgv/datasets/cockpit-ui/raw/debug/ca_car.jpg \
|
||
|
|
--instruction "把主驾温度调到23度" \
|
||
|
|
--output-dir /data/lgv/runs/cockpit-agent/demo_driver_temp_23
|
||
|
|
```
|
||
|
|
|
||
|
|
Focused CLIs are also available:
|
||
|
|
|
||
|
|
```bash
|
||
|
|
python scripts/run_intent.py --instruction "打开内循环"
|
||
|
|
|
||
|
|
python scripts/run_ui_understanding.py \
|
||
|
|
--image /data/lgv/datasets/cockpit-ui/raw/debug/ca_car.jpg \
|
||
|
|
--instruction "后备箱最大开度设置成80%" \
|
||
|
|
--output-dir /data/lgv/runs/cockpit-agent/ui_understanding
|
||
|
|
```
|
||
|
|
|
||
|
|
A successful task output directory contains:
|
||
|
|
|
||
|
|
- `intent.json`
|
||
|
|
- `function_grounding.json`
|
||
|
|
- `roi.jpg`
|
||
|
|
- `ui_state.json`
|
||
|
|
- `plan.json`
|
||
|
|
- `action_grounding.json` when a non-noop action is proposed
|
||
|
|
- `action.json`
|
||
|
|
- `result.jpg`
|
||
|
|
- `run.json`
|
||
|
|
|
||
|
|
The trace uses explicit stage names: `semantic_action`, `function_target`,
|
||
|
|
`function_grounding`, `roi`, `ui_state`, `action_plan`, `action_target`,
|
||
|
|
`action_grounding`, and `mock_action`. Reusing an output directory replaces
|
||
|
|
the pipeline-owned artifacts so a failed run cannot retain stale results from
|
||
|
|
an earlier run.
|
||
|
|
|
||
|
|
If a stage raises an error, the pipeline writes `failure.json` and a partial
|
||
|
|
`run.json` with `success`, `failed_stage`, `error_type`, `error_message`, and
|
||
|
|
`raw_model_output`. It then re-raises the error so callers cannot mistake a
|
||
|
|
failed run for a successful one.
|
||
|
|
|
||
|
|
`action.json` is always a proposal. A tap contains a pixel and repeat count, a
|
||
|
|
drag contains start/end pixels, and a no-op contains its reason. Nothing is
|
||
|
|
executed outside the process.
|
||
|
|
|
||
|
|
## Model Output Handling
|
||
|
|
|
||
|
|
Intent and ROI UI model replies are parsed as JSON and validated against
|
||
|
|
strict dataclass schemas before planning. A small deterministic normalization layer
|
||
|
|
handles known representation-only differences such as `"100%"` versus `100`
|
||
|
|
and Chinese unlock option labels versus `all_doors`; it does not invent missing
|
||
|
|
visual state. UI observation fields are nullable. Slider observations separate
|
||
|
|
endpoint values from `current_value` and expose nullable `track_visible` and
|
||
|
|
`knob_visible` fields; an unobservable current value remains `null` rather than
|
||
|
|
being guessed. The first implementation proposes a drag only when the model
|
||
|
|
explicitly observes `orientation=horizontal`; unknown and vertical sliders fail
|
||
|
|
instead of silently using horizontal geometry. Device state (`current_state`) is separate from visual choice
|
||
|
|
state (`selection_state`). Common choice-state wording such as
|
||
|
|
`active`/`selected` and `inactive`/`unselected` is normalized without changing
|
||
|
|
the model's predicted control type.
|
||
|
|
|
||
|
|
`SemanticTargetBuilder` creates a complete-control/group query before UI
|
||
|
|
understanding and a specific action-control query after planning. The function
|
||
|
|
bbox is padded by the configured ratio and cropped without resizing. ROI action
|
||
|
|
boxes are converted back to source-camera pixels. Slider drags use the visually
|
||
|
|
grounded knob center as the start and the grounded track plus
|
||
|
|
`normalized_target` as the end. The adapter does not repair predictions with
|
||
|
|
image-, vehicle-, color-, position-, or fixed-ROI rules.
|
||
|
|
|
||
|
|
The benchmark integrity policy is documented in
|
||
|
|
[`docs/generalization_rules.md`](docs/generalization_rules.md). In particular,
|
||
|
|
unit tests validate deterministic program logic, while model benchmarks may
|
||
|
|
fail and must report the model's real output.
|
||
|
|
|
||
|
|
## Tests
|
||
|
|
|
||
|
|
Planner tests are deterministic and do not load a model:
|
||
|
|
|
||
|
|
```bash
|
||
|
|
pytest -v
|
||
|
|
```
|
||
|
|
|
||
|
|
The tests also run with the standard library runner:
|
||
|
|
|
||
|
|
```bash
|
||
|
|
python -m unittest discover -s tests -v
|
||
|
|
```
|