Go to file
2026-08-24 16:49:44 +08:00
configs first commit 2026-08-24 16:49:44 +08:00
docs first commit 2026-08-24 16:49:44 +08:00
scripts first commit 2026-08-24 16:49:44 +08:00
src/cockpit_agent first commit 2026-08-24 16:49:44 +08:00
tests first commit 2026-08-24 16:49:44 +08:00
.gitignore first commit 2026-08-24 16:49:44 +08:00
pyproject.toml first commit 2026-08-24 16:49:44 +08:00
README.md first commit 2026-08-24 16:49:44 +08:00

cockpit-agent

cockpit-agent implements a single-pass automotive cockpit task pipeline:

instruction
  -> SemanticAction
  -> function semantic target
  -> function grounding
  -> ROI crop
  -> UIState
  -> ActionPlan
  -> action grounding (when needed)
  -> MockAction

It owns intent parsing, current UI understanding, deterministic planning, and mock action generation. The existing cockpit-ui-grounding package owns:

image + target -> bbox -> pixel (u, v)

The current stage is deliberately limited to NO ROBOT, NO FEEDBACK, and NO CLOSED LOOP. It does not connect to ROS, a mechanical arm, or any real touch API, and it does not verify, retry, or replan an action.

Setup

All model paths in configs/default.toml are local. No network model ID is used.

conda activate qwen3vl
pip install -e /data/lgv/projects/cockpit-ui-grounding
pip install -e /data/lgv/projects/cockpit-agent
export CUDA_VISIBLE_DEVICES=0

The default configuration uses one shared Qwen3.5-2B instance for intent, UI understanding, and grounding. ModelRegistry caches models by backend and resolved local path, so identical settings do not load duplicate weights. Function ROI padding is configured independently with perception.roi_padding_ratio; the default is 0.15.

Run A Task

python scripts/run_task.py \
  --image /data/lgv/datasets/cockpit-ui/raw/debug/ca_car.jpg \
  --instruction "把主驾温度调到23度" \
  --output-dir /data/lgv/runs/cockpit-agent/demo_driver_temp_23

Focused CLIs are also available:

python scripts/run_intent.py --instruction "打开内循环"

python scripts/run_ui_understanding.py \
  --image /data/lgv/datasets/cockpit-ui/raw/debug/ca_car.jpg \
  --instruction "后备箱最大开度设置成80%" \
  --output-dir /data/lgv/runs/cockpit-agent/ui_understanding

A successful task output directory contains:

  • intent.json
  • function_grounding.json
  • roi.jpg
  • ui_state.json
  • plan.json
  • action_grounding.json when a non-noop action is proposed
  • action.json
  • result.jpg
  • run.json

The trace uses explicit stage names: semantic_action, function_target, function_grounding, roi, ui_state, action_plan, action_target, action_grounding, and mock_action. Reusing an output directory replaces the pipeline-owned artifacts so a failed run cannot retain stale results from an earlier run.

If a stage raises an error, the pipeline writes failure.json and a partial run.json with success, failed_stage, error_type, error_message, and raw_model_output. It then re-raises the error so callers cannot mistake a failed run for a successful one.

action.json is always a proposal. A tap contains a pixel and repeat count, a drag contains start/end pixels, and a no-op contains its reason. Nothing is executed outside the process.

Model Output Handling

Intent and ROI UI model replies are parsed as JSON and validated against strict dataclass schemas before planning. A small deterministic normalization layer handles known representation-only differences such as "100%" versus 100 and Chinese unlock option labels versus all_doors; it does not invent missing visual state. UI observation fields are nullable. Slider observations separate endpoint values from current_value and expose nullable track_visible and knob_visible fields; an unobservable current value remains null rather than being guessed. The first implementation proposes a drag only when the model explicitly observes orientation=horizontal; unknown and vertical sliders fail instead of silently using horizontal geometry. Device state (current_state) is separate from visual choice state (selection_state). Common choice-state wording such as active/selected and inactive/unselected is normalized without changing the model's predicted control type.

SemanticTargetBuilder creates a complete-control/group query before UI understanding and a specific action-control query after planning. The function bbox is padded by the configured ratio and cropped without resizing. ROI action boxes are converted back to source-camera pixels. Slider drags use the visually grounded knob center as the start and the grounded track plus normalized_target as the end. The adapter does not repair predictions with image-, vehicle-, color-, position-, or fixed-ROI rules.

The benchmark integrity policy is documented in docs/generalization_rules.md. In particular, unit tests validate deterministic program logic, while model benchmarks may fail and must report the model's real output.

Tests

Planner tests are deterministic and do not load a model:

pytest -v

The tests also run with the standard library runner:

python -m unittest discover -s tests -v