Skip to main content

How to use dsh-read-url

DeepSeek Harness URL reader: fetch any page and return clean main-content text/Markdown. Auto charset (GBK/GB2312/UTF-8/Big5), token-efficient (6000-char cap, cache, offset), zero deps, no API key. 网页一键读全文 → 干净正文 / 结构化 Markdown

This article is auto-derived from indexed fields (wiki / faq / compatibility_json), not freshly AI-generated.

This article is derived from the plugin's already-indexed fields (wiki / faq / compatibility_json / readme), not freshly generated by AI. Source field is noted at the end of each section.

Quick start

DeepSeek Harness URL reader: fetch any page and return clean main-content text/Markdown. Auto charset (GBK/GB2312/UTF-8/Big5), token-efficient (6000-char cap, cache, offset), zero deps, no API key. 网页一键读全文 → 干净正文 / 结构化 Markdown

— source: plugins.ai_summary

Install & verify

dsh plugin --profile web add dsh-read-url

Run the command above in your DSH Web Profile. Then enable the plugin in the plugin list.

— source: plugins.install

Key points

  • Concurrency capped at 4 (avoids rate-limiting); a failing page is isolated ([失败] + reason in the output) and does not affect the others;
  • Reuses every read_url capability and the session cache (encoding, cleaning, SPA rendering, 5-min cache — repeat batches hit the cache).
  • Same-host only; login/API/static-asset paths are skipped; URLs deduped (fragment stripped);
  • Concurrency 2 (gentle on the target site); per-page failures recorded as [失败] without aborting;
  • Output is an indented tree: [depth] title (chars) URL;

— source: plugin_wiki.readme_en (fallback readme_raw)

Pitfalls

Review the upstream repo before installing. This guide is auto-derived from indexed fields and may lag the latest release. If anything contradicts the official docs, treat the upstream source as authoritative.

— source: general rule