Guides & answers

DeepSeek Harness Quickstart

·

Short answer

  • DeepSeek Harness is DeepSeek's own open-source agent harness — a pluginized runtime where "everything is a plugin", with three entry points: the dsh CLI, a Web UI, and a Python SDK. It's a developer preview, and it's the same harness DeepSeek's own benchmark runs use (their "minimal mode" numbers).
  • Fastest path: start the Web UI, paste your DeepSeek API key in Settings → Models (no restart needed), pick a workspace, and send a task. The Python SDK is ~15 lines for the same thing.
  • Two quirks worth knowing before you touch it: keys are write-only through the UI, and hand-entered models are treated as text-only until you declare input: [text, image] — image attachments get refused before they're sent.

We covered how to plug DeepSeek into Claude Code, Codex, GitHub Copilot, and OpenCode in a previous piece — those are external agents using DeepSeek as a backend. DeepSeek Harness is the other direction: DeepSeek shipping its own agent runtime instead of relying on everyone else's. This piece is the walkthrough of that runtime, based on the official docs as they stand on 2026-09-20.

What DeepSeek Harness actually is

The official one-liner: a pluginized SDK for building agent harnesses. In practice that means the harness ships as a small core — model routing, session management, a tool protocol — and everything else a coding agent does (task tools, compaction, skills, workspace prompts) is a plugin you can leave out. The docs' own "minimal" example composition runs with just two model-facing tools and nothing else, which tells you how much of a typical agent harness is optional scaffolding here.

It's worth understanding why this tool matters beyond novelty: DeepSeek's published benchmark results are produced with Harness in its minimal mode (max effort, top_p 0.95, temperature 1.0). When the model vendor's own harness is also the reference harness for their benchmark scores, reproducing their numbers gets a lot less speculative.

The project is a developer preview. The repo lives at github.com/deepseek-ai/deepseek-harness — the README is the install entry point for the dsh CLI, and we're deliberately not paraphrasing its install steps here because we can't currently fetch GitHub to verify them; link, don't guess.

The Web UI in three steps

Start the Web UI from the root README's instructions; the command prints its URL. Then:

  1. Configure a model — open Settings → Models, enter a DeepSeek API key on the DeepSeek card, save. The model route becomes usable immediately, without restarting the server.
  2. Choose a workspace — click Choose workspace and select the directory where you started dsh. The session composer stays unavailable until a workspace is selected; a fresh Web UI has none.
  3. Run a task — e.g. "Summarize this repository and identify its main packages." The agent can read and edit workspace files, run commands, delegate work, and maintain a plan, asking first for operations that require approval under the active permission policy.

One design detail we like: keys are write-only. After saving, the page receives a redacted descriptor, never the literal secret. The key is stored in $DSH_HOME/.credentials.yaml and settings retain only a credential reference. If you've ever had a key flash back at you from a settings page, this is the fix.

Adding models beyond DeepSeek

Harness is provider-agnostic, and the docs split the options three ways:

  • Catalog providers — pick Anthropic, OpenAI, and friends from the installed catalog; the endpoint, protocol, and model list come pre-wired. Just add the key.
  • Native-auth providers — Bedrock, Vertex, Azure, and Codex need their native credentials (AWS credentials + region, ADC project, api-version, OAuth respectively). Filling only the API-key field does not configure these — a quiet trap if you're used to key-only setups.
  • Custom providers — for a company gateway or self-hosted endpoint. You supply a lowercase Provider ID, base URL, protocol, credential, and at least one model.

The Provider ID deserves a warning of its own: it's permanent. Requests, saved sessions, model defaults, and credential references all key off it, so renaming means adding a new provider and deleting the old one. Everything else — display name, base URL, credential — stays editable.

The image-input declaration quirk

Here's the detail most likely to bite someone running DeepSeek behind an OpenAI-compatible proxy: a model you enter by hand is treated as text-only until declared otherwise. Nothing can ask an endpoint which modalities it accepts, so Harness assumes none and refuses image attachments before sending, naming the offending model.

The fix is one line in $DSH_HOME/settings.yaml, because the form has no field for it:

my-gateway:
  models:
    - id: legacy-chat
    - id: vision-preview
      input: [text, image]

input applies per model, so one route can serve both text and vision models. There's also a route-level defaultInput fallback (defaulting to [text]) for models the installed catalog doesn't describe, and a modelOverrides block for catalog providers that have no models list of their own. Note these are claims about your endpoint, not checks — declaring images doesn't make an endpoint serve them.

The Python SDK: same agent, scripted

The SDK is the programmatic alternative to the Web UI, and the docs' example is refreshingly small. After pip install deepseek-harness-sdk (the bundled runtime needs no system Node.js):

from pathlib import Path
from deepseek_harness import DeepSeekHarness

with DeepSeekHarness(
    provider="deepseek-official",
    model="deepseek-v4-flash",
    max_tokens=49_152,
    cwd=str(workspace),
    session_root=str(sessions),
) as harness:
    result = harness.run("Inspect the repository and fix the failing tests.",
                         session_id="example-001")

print(result.final_response)

Two runtime semantics matter here. First, the harness starts its bundled runtime lazily and reuses it until the context manager exits — you're not paying process startup per call. Second, reusing the same session id preserves the session-owned Bash process, including its working directory, exported variables, and shell functions. That makes run() calls with the same id a durable conversation, not a stateless API. Fresh task, fresh id.

Each session drops an uncompressed JSONL log under DSH_SESSION_ROOT containing the assembled model requests and tool calls — which is exactly what you want when debugging why an agent did something.

What the minimal composition actually runs

The checked-in example composition is the clearest statement of Harness's philosophy, and it's worth reading as a table:

PropertyValue
Model resolution--model → DSH_MODEL → deepseek-v4-flash
Model-facing toolsPersistent bash and str_replace_editor only
Bash timeout300 seconds
Editor output limit16,000 characters
Context compactionDisabled
FilesystemBare local backend; absolute paths may reach anything the runtime process can see
Session logUncompressed JSONL under the session root

Everything a "real" agent harness would add — harness identity, workspace prompt text, skills, one-shot Bash, task tools, compaction — is omitted, and sandbox-policy facts are logged as runtime context instead of stuffed into the system prompt. That bare local filesystem line is also your security posture: the minimal composition has no sandbox. Point cwd at something you can afford the runtime to modify.

How it fits with the Claude Code / Codex / Copilot integrations

To keep the two official stories straight: the four agent integrations are about running other tools on DeepSeek's API; Harness is DeepSeek's own agent runtime. They overlap in exactly one spot we've verified: DeepSeek's API natively supports Web Search inside Claude Code — when the model decides a question needs it, it invokes the Web Search tool through DeepSeek's own search API. The cost caveat: each invocation generates additional LLM requests to summarize retrieved content, so "free-looking" searches show up as extra tokens on your bill.

If you want the vendor-neutral framing, our What is MCP primer covers how agent tools and model backends generally connect; Harness is one more entry in that ecosystem, just wearing DeepSeek's badge.

Verdict: for a developer preview, the docs are unusually specific about failure modes — permanent Provider IDs, text-only defaults for hand-entered models, key handling — which is a good sign for how the project is run. If your daily driver is already Claude Code or Codex, Harness is worth a test drive for benchmark-reproduction work in particular; replacing a battle-tested agent is a heavier decision than trying a second one.

All Harness behavior described here was verified against the official documentation site on 2026-09-20. The project is a developer preview — plugin surface, defaults, and configuration keys can change before a stable release.

Sources: DeepSeek Harness Docs (official), deepseek-ai/deepseek-harness on GitHub, DeepSeek API Docs — Claude Code integration