Adds video understanding to plain text DSH agents: frame extraction + optional speech transcription + vision model, outputting structured evidence.
- Language
- JavaScript
- License
- MIT
- Branch
- main
Install
$ dsh plugin --profile web add dsh-video-lensRun the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial
Install via your agent
Install the DeepSeek Harness plugin dundunhan/dsh-video-lens for me: review the repository at https://github.com/dundunhan/dsh-video-lens first, then run the install command and verify the plugin loads successfully.
Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.
One-Sentence Positioning
Give pure text-based DeepSeek Harness agents the ability to "see videos" and "hear videos": extract key frames from local video files, optionally with transcript, send to a vision-capable model, so regular conversational AI can answer "what is this video about / what was said at minute 3".
Core Capabilities
- Use
video_probeto get basic video metadata (container, duration, resolution, frame rate, encoding, audio/subtitle tracks) in one command, directly read by ffprobe - Use
video_analyzefor video content understanding: first use ffmpeg's scene detection filter to find shot boundaries, extract one representative frame per shot, then send all frames (with optional transcript) to a vision LLM, output structured JSON evidence - Use
video_askfor time-anchored Q&A: automatically recognize timestamps in questions (like "at 3:20", "minute 2"), or search for relevant segments in transcript by keyword, then re-extract frames near those time windows to answer, with confidence and source timestamp - Vision model and ASR both use OpenAI-compatible protocol, can point to SiliconFlow, DashScope, self-hosted vLLM/Ollama or any similar endpoint, not locked to any vendor
- All video/audio processing done by system's ffmpeg/ffprobe subprocess, no native decoding in-process, temporary files cleaned up after use
Technical Implementation
- Language: JavaScript (ES Modules,
"type": "module") - Key Dependencies:
@deepseek-ai/dsh-tools(inject and register tools to host),@deepseek-ai/schemastery(declare Config Schema); external binaries ffmpeg/ffprobe - Architecture Pattern: Cordis plugin, mount to host's tool registry via
cordis.patch.yml;apply(ctx, config)registers three tools when receivingtoolsinjection - Entry File:
src/index.js, main entry exportsapply, split by responsibility intoprobe.js/frames.js/vlm.js/asr.js/ask.js
Use Cases
Use when you want DSH agents to answer questions as if they've watched videos—for example, having an Agent summarize a recording, locate when a concept is explained in a course video, or find specific footage from surveillance footage. Text-only conversation models can only read text; this plugin transforms any local video file into "visual+text summary an Agent can understand".
Prerequisites and Compatibility
| Dependency | Min Version | Description |
|---|---|---|
| Node.js | >= 20 | Uses AbortSignal.any, built-in fetch and FormData |
| ffmpeg / ffprobe | 6.0 (recommended) | Must be in PATH; scdet filter from 6.0, older versions auto-degrade to uniform sampling |
| Vision Model Endpoint | OpenAI compatible | Defaults to SiliconFlow Qwen3-VL-8B; replace as needed |
| ASR Endpoint | OpenAI /audio/transcriptions | Optional, won't block vision analysis if not configured |
| DSH | Not declared | README only declares successful integration with dsh-base + dsh-web-app |
Platform: macOS / Linux tested by author; Windows marked as "untested" in README. This plugin has no native module dependencies.
Installation
dsh plugin --profile web add dsh-video-lens
Configuration Options
| Config | Type | Description | Default |
|---|---|---|---|
| visionBaseUrl | string | Vision model endpoint (OpenAI-compatible chat/completions) | https://api.siliconflow.cn/v1 |
| visionModel | string | Vision model name | Qwen/Qwen3-VL-8B-Instruct |
| visionApiKeyEnv | string | Env var name for vision model Key | VIDEO_LENS_API_KEY |
| asrBaseUrl | string | ASR endpoint (OpenAI-compatible /audio/transcriptions) | https://api.siliconflow.cn/v1 |
| asrModel | string | ASR model name | FunAudioLLM/SenseVoiceSmall |
| asrApiKeyEnv | string | Env var name for ASR Key | VIDEO_LENS_ASR_KEY |
| maxFrames | number | Max frames to extract per analysis (adapts to duration) | 12 |
| frameMaxWidth | number | Max frame width, reduces resolution when many frames | 768 |
| frameQuality | number | ffmpeg JPEG quality (lower = higher quality) | 4 |
| sceneThreshold | number | Scene detection threshold 0-100, higher = fewer cuts | 10 |
| askPaddingSec | number | Seconds to add on each side of matched transcript segment for video_ask | 2 |
| vlmMaxTokens | number | Max tokens for vision model single response | 1500 |
| vlmTimeoutMs | number | Vision model request timeout (ms) | 90000 |
| asrTimeoutMs | number | ASR request timeout (ms) | 120000 |
FAQ
Q: Do I need to install ffmpeg?
A: Yes. The plugin invokes system ffprobe and ffmpeg subprocesses, requiring them available. ffmpeg >= 6.0 recommended for scdet filter scene detection; older versions auto-degrade to uniform mid-frame sampling, maintaining functionality but with coarser frame extraction strategy. macOS use brew install ffmpeg, Linux use apt install ffmpeg.
Q: Must I configure two Keys (vision + speech)?
A: Vision Key is required—unset visionApiKeyEnv pointing to an environment variable throws an error; Speech Key is optional. If not configured or service is abnormal, transcript will be null, but the vision analysis path still completes.
Q: Where are video files uploaded?
A: Only sent to endpoints you configure in visionBaseUrl and asrBaseUrl, at most once per video_analyze (vision sends frames, ASR sends audio). Defaults to SiliconFlow; change to your self-hosted or trusted OpenAI-compatible service.
Q: Can I use it on Windows?
A: No platform-specific calls at code level (uses node:child_process.execFile cross-platform), but README explicitly marks Windows as "untested"—author hasn't done compatibility validation.
Q: What if the video has no audio?
A: video_analyze / video_ask first probe with ffprobe whether there's an audio track; if not, skip ASR, transcript field is null, return only vision analysis result.
Q: Can I specify what to focus on?
A: video_analyze accepts an optional question parameter, which gets passed as additional instruction to the vision model, letting it focus analysis on specific aspects.
Q: How does video_ask locate the timestamp corresponding to the question?
A: First identify explicit timestamps in the question (like "at 3:20", "minute 2", "3分20秒"); if none found, fall back to keyword matching on the transcript (mixed Chinese/English supported), expanding matched segments by askPaddingSec seconds on each side as time window, re-extract frames within that window for the model to answer, and return confidence with source timestamp.
Q: How to uninstall?
A: Remove dsh-video-lens dependency from profile's package.json, remove corresponding entry from dsh.profile.bundles, then reinstall profile. Plugin writes no persistent data; temporary files self-clean.
Getting Started Difficulty
Beginner — install ffmpeg, configure two environment variables, and it's ready to use. All config options have sensible defaults; regular users don't need to understand concepts like scdet / threshold to get it running.
Known Issues and Limitations
- Windows platform untested by author (README:140), only macOS / Linux verified
- ffmpeg version below 6.0 doesn't have
scdetfilter; framework auto-degrades to time-equidistant frame extraction, quality degrades for videos with dense long takes - DSH has no official plugin review mechanism; plugins run with trusted process permissions (SECURITY.md:5), read "Threat Model and Mitigation" in SECURITY.md before running
- Videos with filenames starting with
-are rejected (assertSafePathdefends against ffmpeg option injection); callers need to pass valid filenames or rename - Max 3 concurrent ffmpeg subprocesses at the same time (
MAX_CONCURRENCY = 3), expected bottleneck when analyzing long videos - Both vision model and ASR depend on external network; network exceptions return as timeout errors, no retry between vision/speech
Video understanding for DeepSeek Harness — give text-only agents eyes and ears on video.
A DeepSeek Harness (DSH) plugin that lets text-only LLM agents understand local video files. It provides two tools:
| Tool | What it does |
|---|---|
video_probe | Cheap, instant metadata via ffprobe: container, duration, resolution, fps, codecs, audio tracks, subtitles. |
video_analyze | Content understanding: scene-change-aware frame sampling (ffmpeg scdet), optional ASR transcript (speech with timestamps), fused with any OpenAI-compatible vision model into structured evidence JSON. |
video_ask | Time-anchored Q&A: parses explicit time references ("at 3:20", "第2分钟") or locates relevant speech via transcript keyword matching, re-samples frames from the matched windows, and answers with grounded evidence (answer + confidence + supporting timestamps). |
v0.3.1. The plugin never locks you into a provider: vision and ASR are both OpenAI-compatible endpoints configured via
baseUrl+model+ key env var.
How it works
video file ──► video_probe ──► ffprobe ──► compact metadata JSON
└─► video_analyze ──► scdet scene detection ──► shot boundaries
├─► ffmpeg frame sampling (one representative frame per shot, capped)
├─► ffmpeg audio extract ──► ASR transcript (timestamped) [optional]
└─► OpenAI-compatible vision API ──► evidence JSON
- Scene changes are detected with ffmpeg's
scdetfilter (ffmpeg ≥ 6.0). Videos without detectable cuts fall back to uniform midpoint sampling. - ASR is strictly additive: if
asrApiKeyEnvis unset or the provider fails, the visual analysis still completes andtranscriptisnull. - All media work is delegated to
ffmpeg/ffprobeonPATH— no native decoding in the agent.
Install
Prerequisites: Node.js ≥ 20, ffmpeg ≥ 6.0 (recommended) with ffprobe on PATH (brew install ffmpeg / apt install ffmpeg).
Option A — npm (recommended)
# in your DSH profile directory (the one containing package.json)
pnpm add dsh-video-lens
Option B — from source (development)
Clone the repo, then mount it into your DSH profile via a local link:
git clone https://github.com/dundunhan/dsh-video-lens.git
Either way, register the bundle in your profile's package.json — this exact block is the full profile configuration:
{
"dependencies": {
"dsh-video-lens": "^0.3"
},
"dsh": {
"profile": {
"bundles": [
"@deepseek-ai/dsh-base",
"@deepseek-ai/dsh-web-app",
"dsh-video-lens"
]
}
}
}
Then export the keys and restart the profile:
export VIDEO_LENS_API_KEY=sk-... # vision
export VIDEO_LENS_ASR_KEY=sk-... # optional, ASR
Configuration
All options are DSH config values:
| Key | Default | Meaning |
|---|---|---|
visionBaseUrl | https://api.siliconflow.cn/v1 | Vision endpoint (OpenAI-compatible) |
visionModel | Qwen/Qwen3-VL-8B-Instruct | Vision model name |
visionApiKeyEnv | VIDEO_LENS_API_KEY | Env var holding the vision key |
asrBaseUrl | https://api.siliconflow.cn/v1 | ASR endpoint (OpenAI-compatible /audio/transcriptions) |
asrModel | FunAudioLLM/SenseVoiceSmall | ASR model name |
asrApiKeyEnv | VIDEO_LENS_ASR_KEY | Env var holding the ASR key |
maxFrames | 12 | Frame budget cap (1–max); actual count is duration-adaptive (~1 frame per 30s, denser for short videos) |
frameMaxWidth | 768 | Max frame width; keeps payloads small |
frameQuality | 4 | JPEG quality (ffmpeg -q:v) |
sceneThreshold | 10 | scdet threshold (0–100); higher = fewer cuts |
askPaddingSec | 2 | video_ask window padding around matched transcript segments |
vlmMaxTokens | 1500 | Vision model max output tokens |
vlmTimeoutMs | 90000 | Vision call timeout |
asrTimeoutMs | 120000 | ASR call timeout |
Usage
Ask the agent:
"What's in /tmp/demo.mp4?"
The agent calls video_probe first, then video_analyze. Evidence includes:
{
"metadata": { "container": "mov,mp4,m4a,3gp,3g2,mj2", "durationSec": 268.4, "...": "..." },
"shots": [{ "timeSec": 12.3, "score": 45.2 }],
"framesSampled": [{ "timestampSec": 5.5, "jpegBytes": 12345 }],
"transcript": {
"text": "…",
"segments": [{ "start": 0.0, "end": 2.4, "text": "…" }],
"language": "zh"
},
"visionModel": "Qwen/Qwen3-VL-8B-Instruct",
"analysis": { "overall_summary": "…", "timeline": [{"timestamp_sec": 5.5, "description": "…"}], "on_screen_text": "…", "visual_style": "…", "notable_moments": "…" }
}
Permissions & security
Read this before using or redistributing. DSH plugins run in the host process as trusted code and there is no official plugin review — self-review is on the author. See SECURITY.md.
What this plugin does
- Reads: any local file path the agent passes to its tools (via
ffprobe/ffmpeg). - Executes:
ffprobeandffmpegfromPATH(never a shell — argv arrays only). - Network: one outbound call per
video_analyzeto the configuredvisionBaseUrl(frames + vision key), and optionally one toasrBaseUrl(audio + ASR key). - Does not: execute shells, eval code, phone home, auto-update, or read files on its own.
Operator responsibilities
- Keys are only as safe as the endpoints they are sent to — configure only endpoints you trust.
- The real access boundary is the DSH host sandbox; the plugin's readability check is a UX guard, not a security boundary.
- Payload sizes are bounded:
maxFrames× ~100–300 KB (768px JPEG) per analysis call.
Compatibility
- Tested with DSH profile bundles
@deepseek-ai/dsh-base+@deepseek-ai/dsh-web-app. - Node ≥ 20 (uses
AbortSignal.any/ built-infetch/FormData). - ffmpeg ≥ 6.0 for
scdet; older versions degrade to uniform sampling. - macOS / Linux tested; Windows untested.
Uninstall
- Remove
dsh-video-lensfromdsh.profile.bundlesin your profilepackage.json. - Remove the dependency:
pnpm remove dsh-video-lens(npm install) — or delete thelink:entry if you installed from source — then reinstall the profile.
Roadmap
- v1.0: frame caching by file hash, evaluation table in README (5 video types × metrics), publish to npm (in progress).
- Beyond: native video-input models as an optional fast path when the configured VLM supports them.
License
MIT — see LICENSE.
Read the usage guide →
Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.
Listing badge
[](https://deepseek-plugin.org/plugins/dundunhan/dsh-video-lens)Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.