CMVR通用音频与视频智能分析服务
Go to file
lixiaolong b386f003a0 feat(vision): 集成Qwen3.8视觉模型并优化音视频分析功能
- 集成Qwen3.8-27B-FP8快速模型和Qwen3.8-27B精确模型作为SGLang服务
- 添加SGLang API配置选项(SGLANG_FAST_BASE_URL、SGLANG_ACCURATE_BASE_URL等)
- 实现音频分类中的决策聚合算法(topK、nearestWeight、labelMaxDistance)
- 添加视频采样帧限制(VIDEO_SAMPLING_FRAME_LIMIT)和上下文token限制
- 更新健康检查以监控SGLang服务状态
- 实现视频分析的双模式决策策略(快速+精确)
- 添加音频参考文件导入工具(import_audio_references.py)
- 扩展音频分类标签支持FIND_VEHICLE_HORN类别
- 优化视频分析的帧采样策略,始终包含视频尾部帧
- 添加决策策略参数(tuning、decisionPolicy)支持
- 更新配置类以支持新的SGLang和视频参数
- 修改compose配置以支持Qwen3.8模型部署
- 更新音频分类测试用例验证聚合逻辑
- 重构视频测试以支持SGLang API格式和决策策略
2026-08-19 09:28:53 +08:00
app feat(vision): 集成Qwen3.8视觉模型并优化音视频分析功能 2026-08-19 09:28:53 +08:00
data feat(vision): 集成Qwen3.8视觉模型并优化音视频分析功能 2026-08-19 09:28:53 +08:00
deploy feat(vision): 集成Qwen3.8视觉模型并优化音视频分析功能 2026-08-19 09:28:53 +08:00
scripts feat: add media analysis service 2026-08-13 16:56:32 +08:00
tests feat(vision): 集成Qwen3.8视觉模型并优化音视频分析功能 2026-08-19 09:28:53 +08:00
tools feat(vision): 集成Qwen3.8视觉模型并优化音视频分析功能 2026-08-19 09:28:53 +08:00
.dockerignore feat: add media analysis service 2026-08-13 16:56:32 +08:00
.gitattributes feat: add media analysis service 2026-08-13 16:56:32 +08:00
.gitignore feat: add media analysis service 2026-08-13 16:56:32 +08:00
Dockerfile feat: add media analysis service 2026-08-13 16:56:32 +08:00
README.md feat(vision): 集成Qwen3.8视觉模型并优化音视频分析功能 2026-08-19 09:28:53 +08:00
requirements-dev.txt feat: add media analysis service 2026-08-13 16:56:32 +08:00
requirements.txt feat: add media analysis service 2026-08-13 16:56:32 +08:00

CMVR Media Analysis Service

This service provides a project-neutral API for multi-label audio event classification and video analysis. Audio labels are discovered from reference subdirectories, so adding categories does not require code changes.

Runtime layout

The production deployment lives under /data/apps/cmvr-ai-analysis on the model server. Reference media belongs in data/profiles; generated features belong in data/artifacts.

Build an audio profile

Place at least three reference files in each label directory, then run:

docker compose run --rm analysis-service python -m tools.build_audio_profile \
  --profile-code aima.power_state.v1

On the model server the reference directories are:

/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/POWER_ON
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/POWER_OFF
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/ARMED
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/DISARMED
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/FIND_VEHICLE_HORN

After adding or replacing reference files, rebuild the feature library and restart the service profile state:

cd /data/apps/cmvr-ai-analysis
sh scripts/build-audio-profile.sh aima.power_state.v1
docker compose -f deploy/compose.yaml restart analysis-service

Run leave-one-out validation after rebuilding. Each sample is compared only with the other references, so a file cannot obtain a perfect score by matching itself:

sh scripts/validate-audio-profile.sh aima.power_state.v1

To add another sound category later, create a new stable uppercase label directory under references, add its display name to labelNames, upload the reference audio, and rebuild the profile. The classifier and platform workflow component do not need another code change.

Audio profiles can reduce single-sample false positives with decision aggregation:

  • topK selects how many nearest references contribute to each label score.
  • nearestWeight controls the nearest reference weight; the remaining weight is assigned to the mean of the selected references.
  • labelMaxDistance defines stricter distance limits for labels that otherwise have a higher false-positive risk.

API

POST /api/v1/analysis/run accepts requestId, analysisType, profileCode, mediaUrl, options, and context.

The deployed endpoint is http://192.168.28.10:14080. It is called by the platform backend and requires a bearer token.

The platform backend deployment must provide the same token through MEDIA_ANALYSIS_API_KEY. The service token is stored only in /data/apps/cmvr-ai-analysis/deploy/.env; it is not exposed to the browser or workflow JSON.

Video requests may set options.analysisMode to one of:

  • AUTO: use the fast model first and fall back to the accurate model when the result is incomplete.
  • FAST: use the low-latency model only.
  • ACCURATE: use the high-accuracy model only.

Existing workflows without this option are treated as AUTO.

Qwen3.8 video models

The model server runs the video models as separate SGLang services:

  • Qwen/Qwen3.8-27B-FP8 on port 14081 for fast analysis.
  • Qwen/Qwen3.8-27B on port 14082 for accurate analysis.

Both services use a 262K context window and FP8 KV cache. ModelScope downloads are kept under /data/apps/cmvr-ai-analysis/models/modelscope, outside the containers. Start or inspect them with:

cd /data/apps/cmvr-ai-analysis
docker compose -f deploy/compose.sglang.yaml up -d
docker compose -f deploy/compose.sglang.yaml ps

The analysis service calls their OpenAI-compatible APIs. When either SGLang endpoint is unavailable, it falls back to the existing Ollama model for the corresponding mode, provided VISION_OLLAMA_FALLBACK_ENABLED=true.