How to use dsh-scrape-webpage
Reads and extracts content from web pages for DeepSeek Harness, enabling users to import data directly into sessions.
This article is auto-derived from indexed fields (wiki / faq / compatibility_json), not freshly AI-generated.
This article is derived from the plugin's already-indexed fields (wiki / faq / compatibility_json / readme), not freshly generated by AI. Source field is noted at the end of each section.
Quick start
Reads and extracts content from web pages for DeepSeek Harness, enabling users to import data directly into sessions.
— source: plugins.ai_summary
Install & verify
dsh plugin --profile web add dsh-scrape-webpage
Run the command above in your DSH Web Profile. Then enable the plugin in the plugin list.
— source: plugins.install
Key points
- 网页抓取:http/https,自动重定向,非 2xx 状态码优雅返回
- 内容提取:标题、页面描述、正文(截断可调)、H1–H6 标题结构、链接列表(上限 100)
- 自动分析:字符数、词/字数、标题数、链接数、预计阅读时长、语言倾向(中/英/混合)、高频关键词 Top30
- 图片下载:提取
<img src/data-src/srcset>(相对路径自动归一化,过滤图标/logo 等噪音),下载到会话工作区.scrape-images\,返回本地路径供视觉模型read_image分析 - 识图插件接口:发布
scrape.imageAnalyzer服务,识图插件注册分析器后,图片自动交给分析器并把结果附在工具输出中
— source: plugin_wiki.readme_en (fallback readme_raw)
Pitfalls
Review the upstream repo before installing. This guide is auto-derived from indexed fields and may lag the latest release. If anything contradicts the official docs, treat the upstream source as authoritative.
— source: general rule