CMVR-AI-ANALYSIS/README.md
lixiaolong 873a82b0f5 feat(service): 升级AI分析服务支持图像分析和视频详细输出
- 移除Ollama相关配置和回退机制,统一使用SGLang作为视觉模型提供商
- 添加IMAGE_ANALYSIS类型支持,允许对多张图片进行分析
- 实现视频分析的详细输出模式,支持紧凑和详细两种结果格式
- 更新环境变量配置,添加VIDEO_MODEL_FRAME_LIMIT和MAX_IMAGE_BYTES
- 修改compose配置文件中的上下文长度和内存分配参数
- 重构视频采样逻辑,限制单次请求帧数以优化显存使用
- 更新API接口文档,添加mediaUrls参数和详细输出选项说明
- 添加图像分析相关的依赖库opencv-python-headless
- 实现结构化JSON响应格式验证和重试机制
2026-08-21 16:19:35 +08:00

4.7 KiB

CMVR Media Analysis Service

This service provides a project-neutral API for multi-label audio event classification and video analysis. Audio labels are discovered from reference subdirectories, so adding categories does not require code changes.

Runtime layout

The production deployment lives under /data/apps/cmvr-ai-analysis on the model server. Reference media belongs in data/profiles; generated features belong in data/artifacts.

Build an audio profile

Place at least three reference files in each label directory, then run:

docker compose run --rm analysis-service python -m tools.build_audio_profile \
  --profile-code aima.power_state.v1

On the model server the reference directories are:

/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/POWER_ON
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/POWER_OFF
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/ARMED
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/DISARMED
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/FIND_VEHICLE_HORN

After adding or replacing reference files, rebuild the feature library and restart the service profile state:

cd /data/apps/cmvr-ai-analysis
sh scripts/build-audio-profile.sh aima.power_state.v1
docker compose -f deploy/compose.yaml restart analysis-service

Run leave-one-out validation after rebuilding. Each sample is compared only with the other references, so a file cannot obtain a perfect score by matching itself:

sh scripts/validate-audio-profile.sh aima.power_state.v1

To add another sound category later, create a new stable uppercase label directory under references, add its display name to labelNames, upload the reference audio, and rebuild the profile. The classifier and platform workflow component do not need another code change.

Audio profiles can reduce single-sample false positives with decision aggregation:

  • topK selects how many nearest references contribute to each label score.
  • nearestWeight controls the nearest reference weight; the remaining weight is assigned to the mean of the selected references.
  • labelMaxDistance defines stricter distance limits for labels that otherwise have a higher false-positive risk.

API

POST /api/v1/analysis/run accepts requestId, analysisType, profileCode, mediaUrl, mediaUrls, options, and context.

The deployed endpoint is http://192.168.28.10:14080. It is called by the platform backend and requires a bearer token.

The platform backend deployment must provide the same token through MEDIA_ANALYSIS_API_KEY. The service token is stored only in /data/apps/cmvr-ai-analysis/deploy/.env; it is not exposed to the browser or workflow JSON.

Video requests may set options.analysisMode to one of:

  • AUTO: use the fast model first and fall back to the accurate model when the result is incomplete.
  • FAST: use the low-latency model only.
  • ACCURATE: use the high-accuracy model only.

Existing workflows without this option are treated as AUTO.

Video requests may also set options.detailedOutput. It defaults to false and returns a compact verdict without an event timeline. Set it to true only when time-ranged event evidence and warnings are required.

Long videos are sampled uniformly across their full duration. At most 30 frames are sent in one model request so visual encoding stays within the H100 dynamic-memory budget; the first and final state remain represented.

Image requests use analysisType=IMAGE_ANALYSIS and profile common.image_analysis.v1. Supply one to twelve scene image URLs in mediaUrls. Set options.prompt to the complete analysis instruction. The model follows the requested content and format without adding pass/fail fields. The response result contains result (parsed JSON, array, number, or text) and resultText (the exact model text). An optional options.referenceImageUrl is appended as the final image.

Qwen3.8 video models

The model server runs the video models as separate SGLang services:

  • Qwen/Qwen3.8-27B-FP8 on port 14081 for fast analysis.
  • Qwen/Qwen3.8-27B on port 14082 for accurate analysis.

Both services use a 262K context window and FP8 KV cache. ModelScope downloads are kept under /data/apps/cmvr-ai-analysis/models/modelscope, outside the containers. Start or inspect them with:

cd /data/apps/cmvr-ai-analysis
docker compose -f deploy/compose.sglang.yaml up -d
docker compose -f deploy/compose.sglang.yaml ps

The analysis service calls only their OpenAI-compatible APIs. An unavailable SGLang endpoint returns an explicit error; there is no Ollama fallback.