给纯文本 DSH 智能体增加视频理解能力:抽帧 + 可选语音转写 + 视觉模型,输出结构化证据
- 语言
- JavaScript
- License
- MIT
- 分支
- main
安装
$ dsh plugin --profile web add dsh-video-lens在终端中运行以上命令,通过 dsh CLI 安装此插件。可在右上角切换 Profile。 第一次用 dsh?看这篇新手教程
对话式安装
帮我安装 DeepSeek Harness 插件 dundunhan/dsh-video-lens:先查看仓库 https://github.com/dundunhan/dsh-video-lens 确认安全性,然后执行安装命令并验证插件加载成功。
把这段指令粘贴给 DSH Web GUI 里的助手,由它代你完成安装与验证。
一句话定位
给纯文本的 DeepSeek Harness 智能体增加"看视频"和"听视频"的能力:把本地视频文件抽取成若干关键帧,附带可选的语音文字稿,送给一个支持图像理解的模型,让普通对话式 AI 能回答"这个视频讲了什么 / 第 3 分钟说了什么"。
核心能力
- 用
video_probe一行命令拿到视频的基础元信息(容器、时长、分辨率、帧率、编码、是否有音轨和字幕),由 ffprobe 直接读出 - 用
video_analyze对视频做内容理解:先用 ffmpeg 的场景切分滤镜找到镜头边界,每个镜头抽一张代表帧,再把所有帧(连同可选的语音转写稿)交给一个视觉大模型,输出一份结构化 JSON 证据 - 用
video_ask做时间锚定的问答:自动识别问题里的时间戳(如 "at 3:20"、"第 2 分钟"),或在转写稿里按关键词命中相关片段,然后在这些时间窗口附近重新抽帧回答,并给出置信度与依据的时间点 - 视觉模型和语音识别均使用 OpenAI 兼容协议,可指向 SiliconFlow、DashScope、自建 vLLM/Ollama 等任何同类端点,不强制绑定某一家厂商
- 所有视频/音频处理都交给系统的 ffmpeg/ffprobe 子进程,进程内不做原生解码,临时文件用完即删
技术实现
- 语言: JavaScript(ES Modules,
"type": "module") - 关键依赖:
@deepseek-ai/dsh-tools(注入并注册工具到宿主)、@deepseek-ai/schemastery(声明 Config Schema);外部二进制 ffmpeg/ffmpeg - 架构模式: Cordis 插件,通过
cordis.patch.yml将插件挂到 host 的 tool 注册表上;apply(ctx, config)在收到tools注入后调用ctx.tools.register注册三个工具 - 入口文件:
src/index.js,主入口导出apply,按职责拆分到probe.js/frames.js/vlm.js/asr.js/ask.js
适用场景
当你希望 DSH 智能体能像看过视频一样回答问题时使用,例如让 Agent 总结一段录屏、定位一段课程视频中讲解某个概念的时间点、或者从监控录像里找出特定画面。文本对话模型本身只能读字,借助本插件可以把任意本地视频文件变成"Agent 看得懂的图文摘要"。
前置依赖与兼容性
| 依赖 | 最低版本 | 说明 |
|---|---|---|
| Node.js | >= 20 | 使用 AbortSignal.any、内置 fetch 与 FormData |
| ffmpeg / ffprobe | 6.0(推荐) | 必须在 PATH 中;scdet 滤镜来自 6.0,旧版自动退化为均匀采样 |
| 视觉模型端点 | OpenAI 兼容 | 默认指向 SiliconFlow Qwen3-VL-8B;按需替换 |
| 语音识别端点 | OpenAI /audio/transcriptions | 可选,未配置不阻断视觉分析 |
| DSH | 未声明 | README 仅声明已与 dsh-base + dsh-web-app 联调通过 |
平台:macOS / Linux 作者测试通过;Windows 在 README 中标注为"untested"。本插件无原生模块依赖。
安装方式
dsh plugin --profile web add dsh-video-lens
配置项
| 配置 | 类型 | 说明 | 默认值 |
|---|---|---|---|
| visionBaseUrl | string | 视觉模型接入地址(OpenAI 兼容 chat/completions) | https://api.siliconflow.cn/v1 |
| visionModel | string | 视觉模型名称 | Qwen/Qwen3-VL-8B-Instruct |
| visionApiKeyEnv | string | 存放视觉模型 Key 的环境变量名 | VIDEO_LENS_API_KEY |
| asrBaseUrl | string | 语音识别接入地址(OpenAI 兼容 /audio/transcriptions) | https://api.siliconflow.cn/v1 |
| asrModel | string | 语音识别模型名称 | FunAudioLLM/SenseVoiceSmall |
| asrApiKeyEnv | string | 存放 ASR Key 的环境变量名 | VIDEO_LENS_ASR_KEY |
| maxFrames | number | 单次分析最多抽多少帧(实际按时长自适应) | 12 |
| frameMaxWidth | number | 帧图最大宽度,帧多时自动降低分辨率 | 768 |
| frameQuality | number | ffmpeg JPEG 质量(数值越小质量越高) | 4 |
| sceneThreshold | number | 场景切分阈值 0–100,越大切分越少 | 10 |
| askPaddingSec | number | video_ask 在匹配到的转写片段两侧各加多少秒 | 2 |
| vlmMaxTokens | number | 视觉模型单次返回的最大 token 数 | 1500 |
| vlmTimeoutMs | number | 视觉模型请求超时(毫秒) | 90000 |
| asrTimeoutMs | number | 语音识别请求超时(毫秒) | 120000 |
常见问题
Q: 需要安装 ffmpeg 吗?
A: 是的。插件会调用系统的 ffprobe 和 ffmpeg 子进程,需要它们可用,推荐 ffmpeg >= 6.0 以使用 scdet 滤镜进行场景切分;旧版本会自动降级为均匀中间帧采样,不影响功能只是抽帧策略变粗。macOS 用 brew install ffmpeg,Linux 用 apt install ffmpeg。
Q: 必须要配置两个 Key(视觉 + 语音)吗?
A: 视觉 Key 是必需的,未配置 visionApiKeyEnv 指向的环境变量会直接报错;语音 Key 是可选的,未配置或服务异常时 transcript 会是 null,但视觉分析路径仍然完成。
Q: 视频文件会上传到哪些地方?
A: 仅发到你在 visionBaseUrl 和 asrBaseUrl 配置的端点,每次 video_analyze 至多各一次(视觉发帧图、ASR 发音频)。默认指向 SiliconFlow,可改为你自建或信任的任何 OpenAI 兼容服务。
Q: Windows 上能用吗?
A: 代码层面没有平台特定调用(用 node:child_process.execFile 跨平台),但 README 明确标注 Windows 未测试,作者未做兼容性验证。
Q: 视频里没有声音会怎样?
A: video_analyze / video_ask 会先 ffprobe 探测是否有音轨,没有就跳过 ASR,transcript 字段为 null,仅返回视觉部分的结果。
Q: 能指定关注的重点吗?
A: video_analyze 支持一个可选的 question 参数,会作为附加指令传给视觉模型,让它分析时重点关注某方面。
Q: video_ask 是怎么定位问题对应的时间点的?
A: 先在问题里识别显式时间戳(如 "at 3:20"、"第 2 分钟"、"3 分 20 秒");没有就退回到语音转写稿上做关键词匹配(中英文混合),把命中的片段前后各扩 askPaddingSec 秒作为时间窗口,在窗口内重新抽帧给模型回答,并返回置信度。
Q: 怎么卸载?
A: 从 profile 的 package.json 中删除 dsh-video-lens 依赖,并在 dsh.profile.bundles 中移除对应条目,然后重新安装 profile。插件不写任何持久数据,临时文件会自清理。
上手难度
入门 — 装好 ffmpeg、配两个环境变量就能用,配置项都有合理默认,普通用户不需要理解 scdet / scdet 阈值等概念也能跑通。
已知问题与限制
- Windows 平台未经过作者测试(README:140),仅 macOS / Linux 验证通过
- ffmpeg 版本低于 6.0 时
scdet滤镜不可用,框架自动降级为按时间等距抽帧,对长镜头密集的视频理解质量会下降 - DSH 没有官方插件审核机制,插件以可信任进程权限运行(SECURITY.md:5),运行前应阅读 SECURITY.md 中的"威胁模型与缓解"
- 文件名以
-开头的视频会被拒绝(assertSafePath防御 ffmpeg 选项注入),调用方需要传入合法文件名或重命名 - 同一时刻最多并发 3 个 ffmpeg 子进程(
MAX_CONCURRENCY = 3),分析长视频时这是预期瓶颈 - 视觉模型与 ASR 都依赖外部网络;网络异常会以超时错误返回,且视觉/语音之间互不重试
Video understanding for DeepSeek Harness — give text-only agents eyes and ears on video.
A DeepSeek Harness (DSH) plugin that lets text-only LLM agents understand local video files. It provides two tools:
| Tool | What it does |
|---|---|
video_probe | Cheap, instant metadata via ffprobe: container, duration, resolution, fps, codecs, audio tracks, subtitles. |
video_analyze | Content understanding: scene-change-aware frame sampling (ffmpeg scdet), optional ASR transcript (speech with timestamps), fused with any OpenAI-compatible vision model into structured evidence JSON. |
video_ask | Time-anchored Q&A: parses explicit time references ("at 3:20", "第2分钟") or locates relevant speech via transcript keyword matching, re-samples frames from the matched windows, and answers with grounded evidence (answer + confidence + supporting timestamps). |
v0.3.1. The plugin never locks you into a provider: vision and ASR are both OpenAI-compatible endpoints configured via
baseUrl+model+ key env var.
How it works
video file ──► video_probe ──► ffprobe ──► compact metadata JSON
└─► video_analyze ──► scdet scene detection ──► shot boundaries
├─► ffmpeg frame sampling (one representative frame per shot, capped)
├─► ffmpeg audio extract ──► ASR transcript (timestamped) [optional]
└─► OpenAI-compatible vision API ──► evidence JSON
- Scene changes are detected with ffmpeg's
scdetfilter (ffmpeg ≥ 6.0). Videos without detectable cuts fall back to uniform midpoint sampling. - ASR is strictly additive: if
asrApiKeyEnvis unset or the provider fails, the visual analysis still completes andtranscriptisnull. - All media work is delegated to
ffmpeg/ffprobeonPATH— no native decoding in the agent.
Install
Prerequisites: Node.js ≥ 20, ffmpeg ≥ 6.0 (recommended) with ffprobe on PATH (brew install ffmpeg / apt install ffmpeg).
Option A — npm (recommended)
# in your DSH profile directory (the one containing package.json)
pnpm add dsh-video-lens
Option B — from source (development)
Clone the repo, then mount it into your DSH profile via a local link:
git clone https://github.com/dundunhan/dsh-video-lens.git
Either way, register the bundle in your profile's package.json — this exact block is the full profile configuration:
{
"dependencies": {
"dsh-video-lens": "^0.3"
},
"dsh": {
"profile": {
"bundles": [
"@deepseek-ai/dsh-base",
"@deepseek-ai/dsh-web-app",
"dsh-video-lens"
]
}
}
}
Then export the keys and restart the profile:
export VIDEO_LENS_API_KEY=sk-... # vision
export VIDEO_LENS_ASR_KEY=sk-... # optional, ASR
Configuration
All options are DSH config values:
| Key | Default | Meaning |
|---|---|---|
visionBaseUrl | https://api.siliconflow.cn/v1 | Vision endpoint (OpenAI-compatible) |
visionModel | Qwen/Qwen3-VL-8B-Instruct | Vision model name |
visionApiKeyEnv | VIDEO_LENS_API_KEY | Env var holding the vision key |
asrBaseUrl | https://api.siliconflow.cn/v1 | ASR endpoint (OpenAI-compatible /audio/transcriptions) |
asrModel | FunAudioLLM/SenseVoiceSmall | ASR model name |
asrApiKeyEnv | VIDEO_LENS_ASR_KEY | Env var holding the ASR key |
maxFrames | 12 | Frame budget cap (1–max); actual count is duration-adaptive (~1 frame per 30s, denser for short videos) |
frameMaxWidth | 768 | Max frame width; keeps payloads small |
frameQuality | 4 | JPEG quality (ffmpeg -q:v) |
sceneThreshold | 10 | scdet threshold (0–100); higher = fewer cuts |
askPaddingSec | 2 | video_ask window padding around matched transcript segments |
vlmMaxTokens | 1500 | Vision model max output tokens |
vlmTimeoutMs | 90000 | Vision call timeout |
asrTimeoutMs | 120000 | ASR call timeout |
Usage
Ask the agent:
"What's in /tmp/demo.mp4?"
The agent calls video_probe first, then video_analyze. Evidence includes:
{
"metadata": { "container": "mov,mp4,m4a,3gp,3g2,mj2", "durationSec": 268.4, "...": "..." },
"shots": [{ "timeSec": 12.3, "score": 45.2 }],
"framesSampled": [{ "timestampSec": 5.5, "jpegBytes": 12345 }],
"transcript": {
"text": "…",
"segments": [{ "start": 0.0, "end": 2.4, "text": "…" }],
"language": "zh"
},
"visionModel": "Qwen/Qwen3-VL-8B-Instruct",
"analysis": { "overall_summary": "…", "timeline": [{"timestamp_sec": 5.5, "description": "…"}], "on_screen_text": "…", "visual_style": "…", "notable_moments": "…" }
}
Permissions & security
Read this before using or redistributing. DSH plugins run in the host process as trusted code and there is no official plugin review — self-review is on the author. See SECURITY.md.
What this plugin does
- Reads: any local file path the agent passes to its tools (via
ffprobe/ffmpeg). - Executes:
ffprobeandffmpegfromPATH(never a shell — argv arrays only). - Network: one outbound call per
video_analyzeto the configuredvisionBaseUrl(frames + vision key), and optionally one toasrBaseUrl(audio + ASR key). - Does not: execute shells, eval code, phone home, auto-update, or read files on its own.
Operator responsibilities
- Keys are only as safe as the endpoints they are sent to — configure only endpoints you trust.
- The real access boundary is the DSH host sandbox; the plugin's readability check is a UX guard, not a security boundary.
- Payload sizes are bounded:
maxFrames× ~100–300 KB (768px JPEG) per analysis call.
Compatibility
- Tested with DSH profile bundles
@deepseek-ai/dsh-base+@deepseek-ai/dsh-web-app. - Node ≥ 20 (uses
AbortSignal.any/ built-infetch/FormData). - ffmpeg ≥ 6.0 for
scdet; older versions degrade to uniform sampling. - macOS / Linux tested; Windows untested.
Uninstall
- Remove
dsh-video-lensfromdsh.profile.bundlesin your profilepackage.json. - Remove the dependency:
pnpm remove dsh-video-lens(npm install) — or delete thelink:entry if you installed from source — then reinstall the profile.
Roadmap
- v1.0: frame caching by file hash, evaluation table in README (5 video types × metrics), publish to npm (in progress).
- Beyond: native video-input models as an optional fast path when the configured VLM supports them.
License
MIT — see LICENSE.
查看使用指南 →
该插件的安装步骤、关键要点、FAQ 与兼容性说明(基于已收录字段派生)。
收录徽章
[](https://deepseek-plugin.org/plugins/dundunhan/dsh-video-lens)把这段 markdown 粘贴到你的 GitHub README,链接回本插件详情页。徽章只声明已被本站收录,不代表安全认证。