CMVR-AI-ANALYSIS/README.md
lixiaolong 873a82b0f5 feat(service): 升级AI分析服务支持图像分析和视频详细输出
- 移除Ollama相关配置和回退机制,统一使用SGLang作为视觉模型提供商
- 添加IMAGE_ANALYSIS类型支持,允许对多张图片进行分析
- 实现视频分析的详细输出模式,支持紧凑和详细两种结果格式
- 更新环境变量配置,添加VIDEO_MODEL_FRAME_LIMIT和MAX_IMAGE_BYTES
- 修改compose配置文件中的上下文长度和内存分配参数
- 重构视频采样逻辑,限制单次请求帧数以优化显存使用
- 更新API接口文档,添加mediaUrls参数和详细输出选项说明
- 添加图像分析相关的依赖库opencv-python-headless
- 实现结构化JSON响应格式验证和重试机制
2026-08-21 16:19:35 +08:00

106 lines
4.7 KiB
Markdown

# CMVR Media Analysis Service
This service provides a project-neutral API for multi-label audio event classification and video analysis. Audio labels are discovered from reference subdirectories, so adding categories does not require code changes.
## Runtime layout
The production deployment lives under `/data/apps/cmvr-ai-analysis` on the model server. Reference media belongs in `data/profiles`; generated features belong in `data/artifacts`.
## Build an audio profile
Place at least three reference files in each label directory, then run:
```bash
docker compose run --rm analysis-service python -m tools.build_audio_profile \
--profile-code aima.power_state.v1
```
On the model server the reference directories are:
```text
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/POWER_ON
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/POWER_OFF
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/ARMED
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/DISARMED
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/FIND_VEHICLE_HORN
```
After adding or replacing reference files, rebuild the feature library and restart the service profile state:
```bash
cd /data/apps/cmvr-ai-analysis
sh scripts/build-audio-profile.sh aima.power_state.v1
docker compose -f deploy/compose.yaml restart analysis-service
```
Run leave-one-out validation after rebuilding. Each sample is compared only with the
other references, so a file cannot obtain a perfect score by matching itself:
```bash
sh scripts/validate-audio-profile.sh aima.power_state.v1
```
To add another sound category later, create a new stable uppercase label directory
under `references`, add its display name to `labelNames`, upload the reference audio,
and rebuild the profile. The classifier and platform workflow component do not need
another code change.
Audio profiles can reduce single-sample false positives with decision aggregation:
- `topK` selects how many nearest references contribute to each label score.
- `nearestWeight` controls the nearest reference weight; the remaining weight is
assigned to the mean of the selected references.
- `labelMaxDistance` defines stricter distance limits for labels that otherwise have
a higher false-positive risk.
## API
`POST /api/v1/analysis/run` accepts `requestId`, `analysisType`, `profileCode`, `mediaUrl`, `mediaUrls`, `options`, and `context`.
The deployed endpoint is `http://192.168.28.10:14080`. It is called by the platform backend and requires a bearer token.
The platform backend deployment must provide the same token through `MEDIA_ANALYSIS_API_KEY`. The service token is stored only in `/data/apps/cmvr-ai-analysis/deploy/.env`; it is not exposed to the browser or workflow JSON.
Video requests may set `options.analysisMode` to one of:
- `AUTO`: use the fast model first and fall back to the accurate model when the result is incomplete.
- `FAST`: use the low-latency model only.
- `ACCURATE`: use the high-accuracy model only.
Existing workflows without this option are treated as `AUTO`.
Video requests may also set `options.detailedOutput`. It defaults to `false` and
returns a compact verdict without an event timeline. Set it to `true` only when
time-ranged event evidence and warnings are required.
Long videos are sampled uniformly across their full duration. At most 30 frames are
sent in one model request so visual encoding stays within the H100 dynamic-memory
budget; the first and final state remain represented.
Image requests use `analysisType=IMAGE_ANALYSIS` and profile
`common.image_analysis.v1`. Supply one to twelve scene image URLs in `mediaUrls`.
Set `options.prompt` to the complete analysis instruction. The model follows the
requested content and format without adding pass/fail fields. The response result
contains `result` (parsed JSON, array, number, or text) and `resultText` (the exact
model text). An optional `options.referenceImageUrl` is appended as the final image.
## Qwen3.8 video models
The model server runs the video models as separate SGLang services:
- `Qwen/Qwen3.8-27B-FP8` on port `14081` for fast analysis.
- `Qwen/Qwen3.8-27B` on port `14082` for accurate analysis.
Both services use a 262K context window and FP8 KV cache. ModelScope downloads are
kept under `/data/apps/cmvr-ai-analysis/models/modelscope`, outside the containers.
Start or inspect them with:
```bash
cd /data/apps/cmvr-ai-analysis
docker compose -f deploy/compose.sglang.yaml up -d
docker compose -f deploy/compose.sglang.yaml ps
```
The analysis service calls only their OpenAI-compatible APIs. An unavailable SGLang
endpoint returns an explicit error; there is no Ollama fallback.