CMVR-AI-ANALYSIS/README.md

106 lines
4.7 KiB
Markdown
Raw Normal View History

2026-08-13 16:56:32 +08:00
# CMVR Media Analysis Service
This service provides a project-neutral API for multi-label audio event classification and video analysis. Audio labels are discovered from reference subdirectories, so adding categories does not require code changes.
## Runtime layout
The production deployment lives under `/data/apps/cmvr-ai-analysis` on the model server. Reference media belongs in `data/profiles`; generated features belong in `data/artifacts`.
## Build an audio profile
Place at least three reference files in each label directory, then run:
```bash
docker compose run --rm analysis-service python -m tools.build_audio_profile \
--profile-code aima.power_state.v1
```
On the model server the reference directories are:
```text
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/POWER_ON
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/POWER_OFF
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/ARMED
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/DISARMED
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/FIND_VEHICLE_HORN
2026-08-13 16:56:32 +08:00
```
After adding or replacing reference files, rebuild the feature library and restart the service profile state:
```bash
cd /data/apps/cmvr-ai-analysis
sh scripts/build-audio-profile.sh aima.power_state.v1
docker compose -f deploy/compose.yaml restart analysis-service
```
Run leave-one-out validation after rebuilding. Each sample is compared only with the
other references, so a file cannot obtain a perfect score by matching itself:
```bash
sh scripts/validate-audio-profile.sh aima.power_state.v1
```
To add another sound category later, create a new stable uppercase label directory
under `references`, add its display name to `labelNames`, upload the reference audio,
and rebuild the profile. The classifier and platform workflow component do not need
another code change.
Audio profiles can reduce single-sample false positives with decision aggregation:
- `topK` selects how many nearest references contribute to each label score.
- `nearestWeight` controls the nearest reference weight; the remaining weight is
assigned to the mean of the selected references.
- `labelMaxDistance` defines stricter distance limits for labels that otherwise have
a higher false-positive risk.
2026-08-13 16:56:32 +08:00
## API
`POST /api/v1/analysis/run` accepts `requestId`, `analysisType`, `profileCode`, `mediaUrl`, `mediaUrls`, `options`, and `context`.
2026-08-13 16:56:32 +08:00
The deployed endpoint is `http://192.168.28.10:14080`. It is called by the platform backend and requires a bearer token.
The platform backend deployment must provide the same token through `MEDIA_ANALYSIS_API_KEY`. The service token is stored only in `/data/apps/cmvr-ai-analysis/deploy/.env`; it is not exposed to the browser or workflow JSON.
Video requests may set `options.analysisMode` to one of:
- `AUTO`: use the fast model first and fall back to the accurate model when the result is incomplete.
- `FAST`: use the low-latency model only.
- `ACCURATE`: use the high-accuracy model only.
Existing workflows without this option are treated as `AUTO`.
Video requests may also set `options.detailedOutput`. It defaults to `false` and
returns a compact verdict without an event timeline. Set it to `true` only when
time-ranged event evidence and warnings are required.
Long videos are sampled uniformly across their full duration. At most 30 frames are
sent in one model request so visual encoding stays within the H100 dynamic-memory
budget; the first and final state remain represented.
Image requests use `analysisType=IMAGE_ANALYSIS` and profile
`common.image_analysis.v1`. Supply one to twelve scene image URLs in `mediaUrls`.
Set `options.prompt` to the complete analysis instruction. The model follows the
requested content and format without adding pass/fail fields. The response result
contains `result` (parsed JSON, array, number, or text) and `resultText` (the exact
model text). An optional `options.referenceImageUrl` is appended as the final image.
## Qwen3.8 video models
The model server runs the video models as separate SGLang services:
- `Qwen/Qwen3.8-27B-FP8` on port `14081` for fast analysis.
- `Qwen/Qwen3.8-27B` on port `14082` for accurate analysis.
Both services use a 262K context window and FP8 KV cache. ModelScope downloads are
kept under `/data/apps/cmvr-ai-analysis/models/modelscope`, outside the containers.
Start or inspect them with:
```bash
cd /data/apps/cmvr-ai-analysis
docker compose -f deploy/compose.sglang.yaml up -d
docker compose -f deploy/compose.sglang.yaml ps
```
The analysis service calls only their OpenAI-compatible APIs. An unavailable SGLang
endpoint returns an explicit error; there is no Ollama fallback.