4.7 KiB
cockpit-agent
cockpit-agent implements a single-pass automotive cockpit task pipeline:
instruction
-> SemanticAction
-> function semantic target
-> function grounding
-> ROI crop
-> UIState
-> ActionPlan
-> action grounding (when needed)
-> MockAction
It owns intent parsing, current UI understanding, deterministic planning, and
mock action generation. The existing cockpit-ui-grounding package owns:
image + target -> bbox -> pixel (u, v)
The current stage is deliberately limited to NO ROBOT, NO FEEDBACK, and NO CLOSED LOOP. It does not connect to ROS, a mechanical arm, or any real touch API, and it does not verify, retry, or replan an action.
Setup
All model paths in configs/default.toml are local. No network model ID is
used.
conda activate qwen3vl
pip install -e /data/lgv/projects/cockpit-ui-grounding
pip install -e /data/lgv/projects/cockpit-agent
export CUDA_VISIBLE_DEVICES=0
The default configuration uses one shared Qwen3.5-2B instance for intent, UI
understanding, and grounding. ModelRegistry caches models by backend and
resolved local path, so identical settings do not load duplicate weights.
Function ROI padding is configured independently with
perception.roi_padding_ratio; the default is 0.15.
Run A Task
python scripts/run_task.py \
--image /data/lgv/datasets/cockpit-ui/raw/debug/ca_car.jpg \
--instruction "把主驾温度调到23度" \
--output-dir /data/lgv/runs/cockpit-agent/demo_driver_temp_23
Focused CLIs are also available:
python scripts/run_intent.py --instruction "打开内循环"
python scripts/run_ui_understanding.py \
--image /data/lgv/datasets/cockpit-ui/raw/debug/ca_car.jpg \
--instruction "后备箱最大开度设置成80%" \
--output-dir /data/lgv/runs/cockpit-agent/ui_understanding
A successful task output directory contains:
intent.jsonfunction_grounding.jsonroi.jpgui_state.jsonplan.jsonaction_grounding.jsonwhen a non-noop action is proposedaction.jsonresult.jpgrun.json
The trace uses explicit stage names: semantic_action, function_target,
function_grounding, roi, ui_state, action_plan, action_target,
action_grounding, and mock_action. Reusing an output directory replaces
the pipeline-owned artifacts so a failed run cannot retain stale results from
an earlier run.
If a stage raises an error, the pipeline writes failure.json and a partial
run.json with success, failed_stage, error_type, error_message, and
raw_model_output. It then re-raises the error so callers cannot mistake a
failed run for a successful one.
action.json is always a proposal. A tap contains a pixel and repeat count, a
drag contains start/end pixels, and a no-op contains its reason. Nothing is
executed outside the process.
Model Output Handling
Intent and ROI UI model replies are parsed as JSON and validated against
strict dataclass schemas before planning. A small deterministic normalization layer
handles known representation-only differences such as "100%" versus 100
and Chinese unlock option labels versus all_doors; it does not invent missing
visual state. UI observation fields are nullable. Slider observations separate
endpoint values from current_value and expose nullable track_visible and
knob_visible fields; an unobservable current value remains null rather than
being guessed. The first implementation proposes a drag only when the model
explicitly observes orientation=horizontal; unknown and vertical sliders fail
instead of silently using horizontal geometry. Device state (current_state) is separate from visual choice
state (selection_state). Common choice-state wording such as
active/selected and inactive/unselected is normalized without changing
the model's predicted control type.
SemanticTargetBuilder creates a complete-control/group query before UI
understanding and a specific action-control query after planning. The function
bbox is padded by the configured ratio and cropped without resizing. ROI action
boxes are converted back to source-camera pixels. Slider drags use the visually
grounded knob center as the start and the grounded track plus
normalized_target as the end. The adapter does not repair predictions with
image-, vehicle-, color-, position-, or fixed-ROI rules.
The benchmark integrity policy is documented in
docs/generalization_rules.md. In particular,
unit tests validate deterministic program logic, while model benchmarks may
fail and must report the model's real output.
Tests
Planner tests are deterministic and do not load a model:
pytest -v
The tests also run with the standard library runner:
python -m unittest discover -s tests -v