Skip to main content

dsh-video-lens

15Stars2Forks0Issues1Watchers

Adds video understanding to plain text DSH agents: frame extraction + optional speech transcription + vision model, outputting structured evidence.

Evidence5/5methodologySourceInstallMaintenanceDSH versionSecurity scan
Machine-auditedInstall commandRepo verifieddsh-plugin topicLicenseREADMEAI wiki
Language
JavaScript
License
MIT
Branch
main
deepseek-harnessdsh-plugin

Install

cmdweb profile
$ dsh plugin --profile web add dsh-video-lens

Run the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial

Install via your agent

Install the DeepSeek Harness plugin dundunhan/dsh-video-lens for me: review the repository at https://github.com/dundunhan/dsh-video-lens first, then run the install command and verify the plugin loads successfully.

Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.

One-Sentence Positioning

Give pure text-based DeepSeek Harness agents the ability to "see videos" and "hear videos": extract key frames from local video files, optionally with transcript, send to a vision-capable model, so regular conversational AI can answer "what is this video about / what was said at minute 3".

Core Capabilities

  • Use video_probe to get basic video metadata (container, duration, resolution, frame rate, encoding, audio/subtitle tracks) in one command, directly read by ffprobe
  • Use video_analyze for video content understanding: first use ffmpeg's scene detection filter to find shot boundaries, extract one representative frame per shot, then send all frames (with optional transcript) to a vision LLM, output structured JSON evidence
  • Use video_ask for time-anchored Q&A: automatically recognize timestamps in questions (like "at 3:20", "minute 2"), or search for relevant segments in transcript by keyword, then re-extract frames near those time windows to answer, with confidence and source timestamp
  • Vision model and ASR both use OpenAI-compatible protocol, can point to SiliconFlow, DashScope, self-hosted vLLM/Ollama or any similar endpoint, not locked to any vendor
  • All video/audio processing done by system's ffmpeg/ffprobe subprocess, no native decoding in-process, temporary files cleaned up after use

Technical Implementation

  • Language: JavaScript (ES Modules, "type": "module")
  • Key Dependencies: @deepseek-ai/dsh-tools (inject and register tools to host), @deepseek-ai/schemastery (declare Config Schema); external binaries ffmpeg/ffprobe
  • Architecture Pattern: Cordis plugin, mount to host's tool registry via cordis.patch.yml; apply(ctx, config) registers three tools when receiving tools injection
  • Entry File: src/index.js, main entry exports apply, split by responsibility into probe.js / frames.js / vlm.js / asr.js / ask.js

Use Cases

Use when you want DSH agents to answer questions as if they've watched videos—for example, having an Agent summarize a recording, locate when a concept is explained in a course video, or find specific footage from surveillance footage. Text-only conversation models can only read text; this plugin transforms any local video file into "visual+text summary an Agent can understand".

Prerequisites and Compatibility

DependencyMin VersionDescription
Node.js>= 20Uses AbortSignal.any, built-in fetch and FormData
ffmpeg / ffprobe6.0 (recommended)Must be in PATH; scdet filter from 6.0, older versions auto-degrade to uniform sampling
Vision Model EndpointOpenAI compatibleDefaults to SiliconFlow Qwen3-VL-8B; replace as needed
ASR EndpointOpenAI /audio/transcriptionsOptional, won't block vision analysis if not configured
DSHNot declaredREADME only declares successful integration with dsh-base + dsh-web-app

Platform: macOS / Linux tested by author; Windows marked as "untested" in README. This plugin has no native module dependencies.

Installation

dsh plugin --profile web add dsh-video-lens

Configuration Options

ConfigTypeDescriptionDefault
visionBaseUrlstringVision model endpoint (OpenAI-compatible chat/completions)https://api.siliconflow.cn/v1
visionModelstringVision model nameQwen/Qwen3-VL-8B-Instruct
visionApiKeyEnvstringEnv var name for vision model KeyVIDEO_LENS_API_KEY
asrBaseUrlstringASR endpoint (OpenAI-compatible /audio/transcriptions)https://api.siliconflow.cn/v1
asrModelstringASR model nameFunAudioLLM/SenseVoiceSmall
asrApiKeyEnvstringEnv var name for ASR KeyVIDEO_LENS_ASR_KEY
maxFramesnumberMax frames to extract per analysis (adapts to duration)12
frameMaxWidthnumberMax frame width, reduces resolution when many frames768
frameQualitynumberffmpeg JPEG quality (lower = higher quality)4
sceneThresholdnumberScene detection threshold 0-100, higher = fewer cuts10
askPaddingSecnumberSeconds to add on each side of matched transcript segment for video_ask2
vlmMaxTokensnumberMax tokens for vision model single response1500
vlmTimeoutMsnumberVision model request timeout (ms)90000
asrTimeoutMsnumberASR request timeout (ms)120000

FAQ

Q: Do I need to install ffmpeg?

A: Yes. The plugin invokes system ffprobe and ffmpeg subprocesses, requiring them available. ffmpeg >= 6.0 recommended for scdet filter scene detection; older versions auto-degrade to uniform mid-frame sampling, maintaining functionality but with coarser frame extraction strategy. macOS use brew install ffmpeg, Linux use apt install ffmpeg.

Q: Must I configure two Keys (vision + speech)?

A: Vision Key is required—unset visionApiKeyEnv pointing to an environment variable throws an error; Speech Key is optional. If not configured or service is abnormal, transcript will be null, but the vision analysis path still completes.

Q: Where are video files uploaded?

A: Only sent to endpoints you configure in visionBaseUrl and asrBaseUrl, at most once per video_analyze (vision sends frames, ASR sends audio). Defaults to SiliconFlow; change to your self-hosted or trusted OpenAI-compatible service.

Q: Can I use it on Windows?

A: No platform-specific calls at code level (uses node:child_process.execFile cross-platform), but README explicitly marks Windows as "untested"—author hasn't done compatibility validation.

Q: What if the video has no audio?

A: video_analyze / video_ask first probe with ffprobe whether there's an audio track; if not, skip ASR, transcript field is null, return only vision analysis result.

Q: Can I specify what to focus on?

A: video_analyze accepts an optional question parameter, which gets passed as additional instruction to the vision model, letting it focus analysis on specific aspects.

Q: How does video_ask locate the timestamp corresponding to the question?

A: First identify explicit timestamps in the question (like "at 3:20", "minute 2", "3分20秒"); if none found, fall back to keyword matching on the transcript (mixed Chinese/English supported), expanding matched segments by askPaddingSec seconds on each side as time window, re-extract frames within that window for the model to answer, and return confidence with source timestamp.

Q: How to uninstall?

A: Remove dsh-video-lens dependency from profile's package.json, remove corresponding entry from dsh.profile.bundles, then reinstall profile. Plugin writes no persistent data; temporary files self-clean.

Getting Started Difficulty

Beginner — install ffmpeg, configure two environment variables, and it's ready to use. All config options have sensible defaults; regular users don't need to understand concepts like scdet / threshold to get it running.

Known Issues and Limitations

  • Windows platform untested by author (README:140), only macOS / Linux verified
  • ffmpeg version below 6.0 doesn't have scdet filter; framework auto-degrades to time-equidistant frame extraction, quality degrades for videos with dense long takes
  • DSH has no official plugin review mechanism; plugins run with trusted process permissions (SECURITY.md:5), read "Threat Model and Mitigation" in SECURITY.md before running
  • Videos with filenames starting with - are rejected (assertSafePath defends against ffmpeg option injection); callers need to pass valid filenames or rename
  • Max 3 concurrent ffmpeg subprocesses at the same time (MAX_CONCURRENCY = 3), expected bottleneck when analyzing long videos
  • Both vision model and ASR depend on external network; network exceptions return as timeout errors, no retry between vision/speech

Read the usage guide →

Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.

Listing badge

Listed on deepseek-plugin.org
[![Listed on deepseek-plugin.org](https://img.shields.io/badge/listed_on-deepseek--plugin.org-007EC6)](https://deepseek-plugin.org/plugins/dundunhan/dsh-video-lens)

Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.

← Back to plugin directory