CMVR-AI-ANALYSIS/README.md
lixiaolong b386f003a0 feat(vision): 集成Qwen3.8视觉模型并优化音视频分析功能
- 集成Qwen3.8-27B-FP8快速模型和Qwen3.8-27B精确模型作为SGLang服务
- 添加SGLang API配置选项(SGLANG_FAST_BASE_URL、SGLANG_ACCURATE_BASE_URL等)
- 实现音频分类中的决策聚合算法(topK、nearestWeight、labelMaxDistance)
- 添加视频采样帧限制(VIDEO_SAMPLING_FRAME_LIMIT)和上下文token限制
- 更新健康检查以监控SGLang服务状态
- 实现视频分析的双模式决策策略(快速+精确)
- 添加音频参考文件导入工具(import_audio_references.py)
- 扩展音频分类标签支持FIND_VEHICLE_HORN类别
- 优化视频分析的帧采样策略,始终包含视频尾部帧
- 添加决策策略参数(tuning、decisionPolicy)支持
- 更新配置类以支持新的SGLang和视频参数
- 修改compose配置以支持Qwen3.8模型部署
- 更新音频分类测试用例验证聚合逻辑
- 重构视频测试以支持SGLang API格式和决策策略
2026-08-19 09:28:53 +08:00

92 lines
3.8 KiB
Markdown

# CMVR Media Analysis Service
This service provides a project-neutral API for multi-label audio event classification and video analysis. Audio labels are discovered from reference subdirectories, so adding categories does not require code changes.
## Runtime layout
The production deployment lives under `/data/apps/cmvr-ai-analysis` on the model server. Reference media belongs in `data/profiles`; generated features belong in `data/artifacts`.
## Build an audio profile
Place at least three reference files in each label directory, then run:
```bash
docker compose run --rm analysis-service python -m tools.build_audio_profile \
--profile-code aima.power_state.v1
```
On the model server the reference directories are:
```text
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/POWER_ON
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/POWER_OFF
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/ARMED
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/DISARMED
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/FIND_VEHICLE_HORN
```
After adding or replacing reference files, rebuild the feature library and restart the service profile state:
```bash
cd /data/apps/cmvr-ai-analysis
sh scripts/build-audio-profile.sh aima.power_state.v1
docker compose -f deploy/compose.yaml restart analysis-service
```
Run leave-one-out validation after rebuilding. Each sample is compared only with the
other references, so a file cannot obtain a perfect score by matching itself:
```bash
sh scripts/validate-audio-profile.sh aima.power_state.v1
```
To add another sound category later, create a new stable uppercase label directory
under `references`, add its display name to `labelNames`, upload the reference audio,
and rebuild the profile. The classifier and platform workflow component do not need
another code change.
Audio profiles can reduce single-sample false positives with decision aggregation:
- `topK` selects how many nearest references contribute to each label score.
- `nearestWeight` controls the nearest reference weight; the remaining weight is
assigned to the mean of the selected references.
- `labelMaxDistance` defines stricter distance limits for labels that otherwise have
a higher false-positive risk.
## API
`POST /api/v1/analysis/run` accepts `requestId`, `analysisType`, `profileCode`, `mediaUrl`, `options`, and `context`.
The deployed endpoint is `http://192.168.28.10:14080`. It is called by the platform backend and requires a bearer token.
The platform backend deployment must provide the same token through `MEDIA_ANALYSIS_API_KEY`. The service token is stored only in `/data/apps/cmvr-ai-analysis/deploy/.env`; it is not exposed to the browser or workflow JSON.
Video requests may set `options.analysisMode` to one of:
- `AUTO`: use the fast model first and fall back to the accurate model when the result is incomplete.
- `FAST`: use the low-latency model only.
- `ACCURATE`: use the high-accuracy model only.
Existing workflows without this option are treated as `AUTO`.
## Qwen3.8 video models
The model server runs the video models as separate SGLang services:
- `Qwen/Qwen3.8-27B-FP8` on port `14081` for fast analysis.
- `Qwen/Qwen3.8-27B` on port `14082` for accurate analysis.
Both services use a 262K context window and FP8 KV cache. ModelScope downloads are
kept under `/data/apps/cmvr-ai-analysis/models/modelscope`, outside the containers.
Start or inspect them with:
```bash
cd /data/apps/cmvr-ai-analysis
docker compose -f deploy/compose.sglang.yaml up -d
docker compose -f deploy/compose.sglang.yaml ps
```
The analysis service calls their OpenAI-compatible APIs. When either SGLang endpoint
is unavailable, it falls back to the existing Ollama model for the corresponding
mode, provided `VISION_OLLAMA_FALLBACK_ENABLED=true`.