- 移除Ollama相关配置和回退机制,统一使用SGLang作为视觉模型提供商 - 添加IMAGE_ANALYSIS类型支持,允许对多张图片进行分析 - 实现视频分析的详细输出模式,支持紧凑和详细两种结果格式 - 更新环境变量配置,添加VIDEO_MODEL_FRAME_LIMIT和MAX_IMAGE_BYTES - 修改compose配置文件中的上下文长度和内存分配参数 - 重构视频采样逻辑,限制单次请求帧数以优化显存使用 - 更新API接口文档,添加mediaUrls参数和详细输出选项说明 - 添加图像分析相关的依赖库opencv-python-headless - 实现结构化JSON响应格式验证和重试机制
106 lines
4.7 KiB
Markdown
106 lines
4.7 KiB
Markdown
# CMVR Media Analysis Service
|
|
|
|
This service provides a project-neutral API for multi-label audio event classification and video analysis. Audio labels are discovered from reference subdirectories, so adding categories does not require code changes.
|
|
|
|
## Runtime layout
|
|
|
|
The production deployment lives under `/data/apps/cmvr-ai-analysis` on the model server. Reference media belongs in `data/profiles`; generated features belong in `data/artifacts`.
|
|
|
|
## Build an audio profile
|
|
|
|
Place at least three reference files in each label directory, then run:
|
|
|
|
```bash
|
|
docker compose run --rm analysis-service python -m tools.build_audio_profile \
|
|
--profile-code aima.power_state.v1
|
|
```
|
|
|
|
On the model server the reference directories are:
|
|
|
|
```text
|
|
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/POWER_ON
|
|
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/POWER_OFF
|
|
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/ARMED
|
|
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/DISARMED
|
|
/data/apps/cmvr-ai-analysis/data/profiles/aima/power-state/v1/references/FIND_VEHICLE_HORN
|
|
```
|
|
|
|
After adding or replacing reference files, rebuild the feature library and restart the service profile state:
|
|
|
|
```bash
|
|
cd /data/apps/cmvr-ai-analysis
|
|
sh scripts/build-audio-profile.sh aima.power_state.v1
|
|
docker compose -f deploy/compose.yaml restart analysis-service
|
|
```
|
|
|
|
Run leave-one-out validation after rebuilding. Each sample is compared only with the
|
|
other references, so a file cannot obtain a perfect score by matching itself:
|
|
|
|
```bash
|
|
sh scripts/validate-audio-profile.sh aima.power_state.v1
|
|
```
|
|
|
|
To add another sound category later, create a new stable uppercase label directory
|
|
under `references`, add its display name to `labelNames`, upload the reference audio,
|
|
and rebuild the profile. The classifier and platform workflow component do not need
|
|
another code change.
|
|
|
|
Audio profiles can reduce single-sample false positives with decision aggregation:
|
|
|
|
- `topK` selects how many nearest references contribute to each label score.
|
|
- `nearestWeight` controls the nearest reference weight; the remaining weight is
|
|
assigned to the mean of the selected references.
|
|
- `labelMaxDistance` defines stricter distance limits for labels that otherwise have
|
|
a higher false-positive risk.
|
|
|
|
## API
|
|
|
|
`POST /api/v1/analysis/run` accepts `requestId`, `analysisType`, `profileCode`, `mediaUrl`, `mediaUrls`, `options`, and `context`.
|
|
|
|
The deployed endpoint is `http://192.168.28.10:14080`. It is called by the platform backend and requires a bearer token.
|
|
|
|
The platform backend deployment must provide the same token through `MEDIA_ANALYSIS_API_KEY`. The service token is stored only in `/data/apps/cmvr-ai-analysis/deploy/.env`; it is not exposed to the browser or workflow JSON.
|
|
|
|
Video requests may set `options.analysisMode` to one of:
|
|
|
|
- `AUTO`: use the fast model first and fall back to the accurate model when the result is incomplete.
|
|
- `FAST`: use the low-latency model only.
|
|
- `ACCURATE`: use the high-accuracy model only.
|
|
|
|
Existing workflows without this option are treated as `AUTO`.
|
|
|
|
Video requests may also set `options.detailedOutput`. It defaults to `false` and
|
|
returns a compact verdict without an event timeline. Set it to `true` only when
|
|
time-ranged event evidence and warnings are required.
|
|
|
|
Long videos are sampled uniformly across their full duration. At most 30 frames are
|
|
sent in one model request so visual encoding stays within the H100 dynamic-memory
|
|
budget; the first and final state remain represented.
|
|
|
|
Image requests use `analysisType=IMAGE_ANALYSIS` and profile
|
|
`common.image_analysis.v1`. Supply one to twelve scene image URLs in `mediaUrls`.
|
|
Set `options.prompt` to the complete analysis instruction. The model follows the
|
|
requested content and format without adding pass/fail fields. The response result
|
|
contains `result` (parsed JSON, array, number, or text) and `resultText` (the exact
|
|
model text). An optional `options.referenceImageUrl` is appended as the final image.
|
|
|
|
## Qwen3.8 video models
|
|
|
|
The model server runs the video models as separate SGLang services:
|
|
|
|
- `Qwen/Qwen3.8-27B-FP8` on port `14081` for fast analysis.
|
|
- `Qwen/Qwen3.8-27B` on port `14082` for accurate analysis.
|
|
|
|
Both services use a 262K context window and FP8 KV cache. ModelScope downloads are
|
|
kept under `/data/apps/cmvr-ai-analysis/models/modelscope`, outside the containers.
|
|
Start or inspect them with:
|
|
|
|
```bash
|
|
cd /data/apps/cmvr-ai-analysis
|
|
docker compose -f deploy/compose.sglang.yaml up -d
|
|
docker compose -f deploy/compose.sglang.yaml ps
|
|
```
|
|
|
|
The analysis service calls only their OpenAI-compatible APIs. An unavailable SGLang
|
|
endpoint returns an explicit error; there is no Ollama fallback.
|