dsh-netguard

Configuration

← dsh-netguard docs

- id: dsh-netguard
  name: 'dsh-netguard'
  config:
    mode: audit                          # or enforce
    allow: ['**.github.com', '*.example.com', 'registry.npmjs.org:443']
    deny: ['*.internal.example']
    allowPrivateAddresses: []            # CIDR blocks; see below
    allowWideWildcards: []               # namespaces you have reviewed; see below
    wideWildcards: warn                  # or refuse, to fail the load on an unreviewed one
    policyFile: ./.dsh-netguard.yml      # optional, lowest trust
    spoolPath: /var/log/dsh/netguard.ocsf.jsonl    # absolute; the host memory sits beside it
    hostMemoryPath: /var/log/dsh/netguard.hosts    # absolute; default <spoolPath>.hosts
    hostMemoryFlushMs: 5000                        # 0 writes the host memory on every decision
    fetchProviderId: dsh-netguard        # what web.fetchProvider has to name
    searchProviderId: dsh-netguard       # what web.searchProvider would have to name
    vendorName: dsh-security-plugins     # metadata.product.vendor_name
    extension:
      name: dsh                          # keys the extension-owned attributes object
      placement: unmapped                # or `attribute`, which puts it at the top level
      uid: 999                           # omit until the OCSF registry assigns you one
    fetch:
      enabled: true
      timeoutMs: 30000
      maxRedirects: 5
      maxResponseBytes: 5000000
      maxBodyChars: 100000
      maxUrlLength: 2048
      # default: dsh-netguard/<this package's version> (+the repository URL)
      userAgent: 'acme-agent/2.0'
    search:
      enabled: true
      maxQueryLength: 2048               # past this a query is denied unscanned
      delegate:                          # absent = the search provider stays unusable
        module: '@deepseek-ai/dsh-web-search-exa'
        export: 'ExaSearchProvider'
        options: { apiKey: '...', baseURL: 'https://api.exa.ai', searchType: auto, highlightsPerResult: 1 }
    shell:
      enabled: true                      # false turns the command-text arm off entirely
      maxCommandLength: 65536            # past this a command is denied unscanned
      readTextHosts: false               # also read hosts written without a scheme
      commands:                          # tool -> the argument, or arguments, to read
        bash: command
        pwsh: command
        run_code: code
        plugin_manager: [target, registry]
    alerts:
      distinctUrlsPerHost: 32            # per session, per host; 0 turns the signal off
    hmacKey: { source: ephemeral }       # or { source: env, variable: NETGUARD_KEY }
    fleet:
      tenantUid: acme
      labels: [prod]
      tags: { team: security }           # metadata.tags[]
      installUid: laptop-7               # skips the sidecar entirely when you set it
      installUidPath: /var/log/dsh/netguard.install-uid   # absolute; default $DSH_HOME/install-uid

Every path is required to be absolute. A relative one resolves against the process’s working directory, which for dsh is the workspace — the same directory the repo-local policy tier is defended against — so a relative spoolPath puts the audit trail somewhere the agent it records can rewrite. A relative path fails the mount.

Two sidecar files are created on first use:

File Holds Matters because
<spoolPath>.hosts every host seen, with first/last sighting and counts is_alert on a first-seen host, and report --suggest
$DSH_HOME/install-uid one minted UUID device.uid, which is stable across a rename and unique across a fleet imaged from one template

The uid sits under the harness home rather than beside the spool because dsh-ocsf-forwarder spools elsewhere and reads the same file: one machine has to report one device.uid, or every SOC query that groups by device splits this host in two. A uid a release up to 0.1.0 left at <spoolPath>.install-uid is read on first run and written through to the new path, so upgrading does not re-identify the host.

Point logrotate at the spool only. The two sidecars are rewritten in place rather than appended to, and rotating them costs the installation its host memory and its device.uid:

/var/log/dsh/netguard.ocsf.jsonl {
  weekly
  rotate 8
  compress
  missingok
  notifempty
  copytruncate
}

The host memory holds one entry per distinct host — about 155 bytes, so 5,000 hosts is a 784 KB file — and the whole of it is rewritten per write, on the agent’s own event loop. The write is therefore debounced: hostMemoryFlushMs bounds how long a sighting sits unwritten, and the memory is flushed when the plugin unloads, which the harness does on ordinary completion and on SIGINT and SIGTERM alike. What a kill that runs no handler costs is one interval of sightings — a host first contacted inside it is reported first-seen once more, and its counts resume from the last write. Set 0 to write on every decision and pay about 6 ms per decision at 5,000 hosts; nothing here changes a verdict either way.

Set fleet.installUid yourself and the uid sidecar is never written. A harness home this process cannot write is reported on stderr and the logger and then continues with an in-memory uid, because losing a stable device.uid is a smaller loss than refusing to mount.

The pattern grammar

Codex’s semantics, which are the only unambiguous ones in the prior art:

Pattern Matches
example.com that host, and nothing else
*.example.com subdomains only — never the apex
**.example.com the apex and every subdomain
* everything; accepted in allow only
example.com:8443 that host on that port only
[::1], [::1]:443 an IPv6 literal, always bracketed
example.com/org/repo that path and everything under it — allow list only

A deny match wins over every allow match, across every configuration source. An empty allow list denies everything, and that is what ships.

A pattern that could be read two ways is refused at load rather than widened. Refused: a prefix wildcard (prod*.blob.core.windows.net — that namespace is self-service, so the pattern matches names an attacker can register), a wildcard anywhere but at the front (a*b.example.com, *.*.internal.example, *.internal.*), a wildcard over a top-level domain (*.com), an unbracketed IPv6 literal, a URL, or credentials. Only a leading *. or **. is a wildcard; anything else is a load-time error, in a deny list as much as in an allow list.

Wildcards over a namespace anyone can register under

*.s3.amazonaws.com reads like a grant to one vendor and covers every bucket anyone can create. An allow entry whose wildcard base is a namespace where a third party registers the name is wide, and netguard says so at load. *.pages.dev, *.workers.dev, *.vercel.app, *.githubusercontent.com, *.github.io, *.blob.core.windows.net, *.herokuapp.com, *.netlify.app are all wide, and so is every name above one of them: *.amazonaws.com is the wider spelling of the same grant, so it counts too. 8,885 namespaces are listed and 9,286 names are wide once those ancestors are counted.

A country registry’s own second level is wide for the same reason — co.ke is where Kenya’s registry starts handing out names, not a company. *.co.uk, *.co.ke, *.com.pk, *.com.ng, *.gob.mx, *.oslo.no and *.k12.ak.us are wide; *.mycompany.co.ke and the exact mycompany.co.ke are not.

Two things stay accepted with nothing said about them, because they are grants to a name you hold:

allow:
  - '*.mybucket.s3.amazonaws.com'   # the names under one bucket
  - 'mybucket.s3.amazonaws.com'     # that host, exactly

A deny entry is not checked. deny: ['*.s3.amazonaws.com'] compiles silently: a deny wider than its author meant refuses more, which is the safe direction.

What netguard does about a wide entry

wideWildcards chooses, for every namespace allowWideWildcards does not name:

wideWildcards A wide allow entry
warn (default) compiles; one advisory per entry goes to ctx.logger.warn and to stderr at every boot
refuse fails the load, quoting the narrower entry and the allowWideWildcards line

The advisory names the namespace, why it is one, what else the entry admits, the Public Suffix List revision that judgement came from, and the line that stops it (wrapped here for the page; the real message puts each labelled part on one line):

dsh-netguard: allow entry '*.s3.amazonaws.com' is a wildcard over "s3.amazonaws.com", a namespace
anyone can register a name under. It is permitted and every decision it clears is recorded as one
a wide entry cleared.
  Why: the Public Suffix List's private section carries "s3.amazonaws.com", which is its owner
declaring that the name one label under it belongs to whoever registered it.
  From: Public Suffix List VERSION 2026-09-03_19-51-30_UTC, retrieved 2026-09-05.
  Narrower: '*.<your-name>.s3.amazonaws.com', or '<your-name>.s3.amazonaws.com' for that one host.
  To record this namespace as reviewed and stop this message, add it to the dsh-netguard row of
the profile's cordis.yml or cordis.patch.yml:

  - id: dsh-netguard
    config:
      allowWideWildcards: ['s3.amazonaws.com']

  To fail the load on an entry like this instead, set wideWildcards: 'refuse' on the same row.

Whether or not the advisory is silenced, every decision a wide entry clears carries the namespace into the record as wide_wildcard, and dsh-netguard report counts them:

permitted by a wide wildcard entry: 314
  a wildcard over a namespace anyone can register a name under; see docs/enforcement.md

by wide wildcard namespace
     314  s3.amazonaws.com

An entry no request used produces no record. The boot advisory is what states the posture then, which is why it is printed at every boot rather than once.

allowWideWildcards: this namespace has been reviewed

allow: ['*.s3.amazonaws.com']
allowWideWildcards: ['s3.amazonaws.com']

One list, one meaning under either posture: naming a namespace here says you looked at what a wildcard over it admits and accepted it. Under refuse that review is what makes the entry compile at all; under warn the entry compiles either way and the review is what stops the advisory. It changes nothing else — the entry is still recorded as wide on every decision it clears.

An entry is the namespace, not a pattern: one line covers both the *. and the **. spelling, and writing '*.s3.amazonaws.com' there is itself an error. It names that namespace only. A top-level domain cannot be named at all, because * in the allow list is already the entry that means every host.

Where the namespaces come from, and when they go stale

The whole Public Suffix List, both sections, is vendored in the package: 6,949 ICANN rules and 3,372 private-section rules, of which the 8,865 with two labels or more are carried. The 287 wildcard rules and 8 exception rules are kept as their own tables too, because the grammar reads them differently: a wildcard rule makes every name one label under its parent wide, and an exception rule takes one name back out of that. The 1,448 one-label rules are dropped, because a wildcard over a top-level domain is refused before this data is read. Nothing fetches any of it — not at build, not at load, not at request time. src/suffixes.ts records the VERSION and COMMIT of the file it was generated from, and scripts/regenerate-suffixes.mjs is how a maintainer moves it.

That data is therefore as old as the release you installed, and it fails in two directions:

The list is also populated by request, so a namespace whose owner never asked is absent from it however freely it hands out hostnames. netguard carries its own table of 20 such names — wordpress.com, slack.com, atlassian.net and the rest — under the rule in limitations.

Path-scoped allow entries

github.com/your-org/your-repo is a meaningfully narrower grant than all of github.com, and the widest entry you can write is exactly what you should not have to. An allow entry may therefore carry a path:

allow:
  - 'huggingface.co/your-org/your-model'
  - 'github.com/your-org/your-repo'

The rules are as narrow as the host grammar, for the same reason — an entry that could be read two ways is refused at load:

   
What it grants the path itself and everything under it, ending on a segment boundary: example.com/api covers /api and /api/v2, never /apiv2
Case case-sensitive, because only a URL’s scheme and host are not. example.com/Org does not cover /org
Traversal . and .. are resolved by the URL parser before the decision, so a path cannot be climbed out of; a request percent-encoding a slash (..%2f) matches no path-scoped entry, because the origin server may split the segments differently than this package does
Query strings not part of the grant. The decision is on the path alone, and a pattern carrying ? is refused rather than silently ignored
Redirects re-decided per hop, so a granted path cannot be used as an open redirector into one that is not
Deny entries host-only. A path on a deny entry would refuse less than the same line without one, so it is a load-time error

Also refused: a trailing slash (example.com/api/ — one grant, one spelling), a wildcard inside the path (example.com/org/*), a percent-encoded slash, a . or .. segment (a request path never carries one, so the entry would match nothing), a path on the bare *, and a path on an IP address — 10.0.0.0/8 is a CIDR block in the field one above, and reading it as a path would be a different policy from the one you wrote.

Two things a path-scoped entry does not do. A request to an allowed host outside its granted path is refused as blocked-by-allowlist, the same reason as a host that is not listed at all — the model is told to ask for the entry it needs, and the path never appears in the message or in a record. And a search query that merely names the host is not refused: the query filter decides whether the host is one this policy tolerates, and it has no path to decide. A result URL from a search is decided against the path, because the model can hand it straight to web_fetch. report --suggest is unaffected: it writes hosts, never paths.

Two tables back these refusals, both shipped inside the package and versioned with it. Nothing is fetched, at install, at build or at run.

A selection from the Public Suffix List. The private-section blocks for the platforms an agent deployment plausibly reaches — public cloud, object storage, CDN, serverless and PaaS, code and model hosting, static-site and preview hosting, developer tunnels, and the large dynamic-DNS providers — and every two-label country-code suffix whose second level is a generic administrative namespace (com., co., org., net., ac., gov., gob., ne., or., sch. and about eighty more).

netguard’s own table, for what that list cannot carry. The Public Suffix List is populated by request, so a namespace whose owner never asked is absent from it however freely it hands out hostnames — wordpress.com, zendesk.com, atlassian.net and slack.com are all absent from both of its sections. Twenty such namespaces are carried here, each admitted by one rule you can run yourself: a label nobody has claimed resolves under the name, and the name’s own operator answers that hostname saying no tenant holds it. A wildcard catch-all that serves the operator’s marketing page fails the second half and is not carried. Every entry records the status and the quoted answer it was admitted on.

A namespace neither table carries is accepted, and is your risk. Left out on purpose: a registry’s geographic second levels (oslo.no, aichi.jp, ny.us) and the structure under them, and the regional web hosts, blog builders and vanity namespaces at the tail of the private section. Measured, not estimated: 1,827 of the list’s 3,300 private-section rules and 4,168 of its 5,501 multi-label ICANN bases still compile as an allow wildcard. Close any of these with a deny entry.

Hosts are compared canonically on both sides. 2130706433, 0x7f000001, 127.1 and 017700000001 are all 127.0.0.1 because the hostname is read from url.hostname and WHATWG URL has already normalised them; [::ffff:127.0.0.1], the [::ffff:7f00:1] spelling URL leaves behind, the deprecated [::127.0.0.1] form and the NAT64 prefix [64:ff9b::7f00:1] are all unwrapped to 127.0.0.1 by this package. A name is IDNA-normalised and a trailing dot is dropped.

The command-text arm

bash, pwsh and run_code reach the network through a child process or a worker thread. No seam, event or registry in the harness observes a byte either of those sends — ctx.tools.guard() and the tools/pre-execute waterfall both run before the spawn, and both are handed the model-authored argument and nothing else. So this arm decides the destinations a command writes out, one policy decision and one record per destination, before the command runs.

plugin_manager is in the same map for the same reason, and it is not a shell. Its install_bundle action says where a package comes from in two model-authored arguments — registry, which reaches pnpm as --registry=<url>, and target, an install spec that may itself be a tarball URL or a repository URL — and the package fetch is a direct execa inside the plugin-manager package rather than a routed shell call, so the tool call is the only point at which any mounted plugin sees the destination named. An install_bundle that names neither argument leaves the destination to pnpm’s own configuration, which this arm cannot see and does not pretend to decide.

Spelling Read by default Decided against
curl https://example.com/a/b yes the host, the port, and the path
curl http://169.254.169.254/ yes the host policy, then the refused-address table
curl example.com only under readTextHosts the host and port alone
git clone git@github.com:org/repo only under readTextHosts the host and port alone
curl "$(cat url.txt)" never — it names no destination nothing

readTextHosts is off because the bare-token heuristic the search-query filter uses reads main.cc, Makefile.in, Makefile.am and socket.io as hosts, and in enforce mode each of those refuses an ordinary build command. Turn it on when the agent’s commands are narrow enough that a false positive is cheaper than a missed curl example.com.

The arm resolves no names: a URL naming an address literal is checked against the refused-address table, and a URL naming a host is not, because a lookup inside a synchronous guard is I/O the agent loop waits on and is itself a request leaving the host for a command that has not run.

commands replaces the built-in map rather than adding to it, so a deployment that mounts a shell tool under another name restates the four built-ins beside it. An empty map means the built-ins; enabled: false is how the arm is turned off. A value may be one argument name or a list of them, read in the order given; a tool named with an empty string, a blank name, or an empty list fails the mount.

Only a host a URL named enters the host memory, so report --suggest never proposes an allow entry for a word that happened to look like a hostname.

The distinct-URL signal

A host allowlist cannot see an exfiltration that only ever contacts an allowed host. The shape CVE-2026-54316 uses is a covert storage channel (CWE-515): with huggingface.co allowlisted as a bare hostname, the secret is carried by which of many URLs on that one host is requested, and it is read back out of the vendor’s own download counters. No response body is needed and no refused host is ever named.

What is visible is the request count. Every full URL is already reduced to an HMAC digest, so this package counts the distinct URL digests per session, per host, writes the count into each record as distinct_urls, and sets is_alert once it reaches alerts.distinctUrlsPerHost. No verdict changes: this is a signal, in the same place first_seen_host is one.

The default of 32 is a judgement, not a measurement. It is well past an agent reading documentation or a handful of a repository’s files in one session, and inside the range a channel carrying even a short secret needs — one request per byte puts a 32-byte token at exactly 32 requests. A patient exfiltrator defeats it: 20 URLs per session, or a channel split across several allowed hosts, never reaches the threshold. Set it lower to catch more and page more often, or to 0 to turn the alert off and keep the count.

Only the URL the model asked for is counted. A redirect target is the server’s choice, so a redirecting host cannot raise the alert on the agents that visit it. The counter holds digests, never URLs, and it is capped the way the tool-call join is: 64 session-and-host pairs, 256 distinct URLs each, so a long session cannot grow it without limit — past those caps the count saturates and the oldest pair is dropped.

Addresses that are never reachable

Every resolved address is checked against a fixed table before the socket opens, and one refused address refuses the whole answer — a name with a public A record and an internal AAAA record reaches the internal host on any client that prefers IPv6, and picking the “good” one would make the outcome depend on address selection order rather than on policy.

Refused: 0.0.0.0/8, 10/8, 100.64/10, 127/8, 169.254/16, 172.16/12, 192.0.0/24, 192.168/16, 198.18/15, 224/4, 240/4, ::/128, ::1/128, fc00::/7, fe80::/10, ff00::/8, and the cloud metadata endpoints 169.254.169.254, 169.254.170.2, 168.63.129.16, fd00:ec2::254.

allowPrivateAddresses opens named blocks for a deployment that genuinely needs an internal service — ['10.0.0.0/8'] for a corporate wiki, ['127.0.0.1/32'] for a local fixture. harden-runner allowlists RFC1918 by default; for an agent on a developer’s own machine or a build host that is the wrong call, so nothing is reachable here unless you name it. The cloud metadata endpoints and the whole link-local range cannot be opened at all: an entry that overlaps them is a load-time error, because an agent that can reach 169.254.169.254 holds the host’s cloud role.

Overlap is tested in both directions, which rules out an entry a deployment is likely to reach for: fd00:ec2::254/128 sits inside fc00::/7, so IPv6 ULA cannot be opened wholesale — allowPrivateAddresses: ['fc00::/7'] fails the mount, and so does ['fd00::/8']. Name the prefix the service actually sits on instead and it loads: ['fd12:3456:789a::/48'] does. The same holds for any block wide enough to contain an absolute entry: 169.0.0.0/8 contains the link-local range and two metadata endpoints, and 0.0.0.0/0 and ::/0 contain everything. That is the rule working — a block that wide opens the metadata endpoint along with whatever you meant by it — and the mount error names the block you overlapped, so the entry to write instead is a narrower one.

Configuration trust ranking

Rank Source May
1 invariants compiled into the package everything; not configurable
2 cordis.yml / bundle patch config set every field, including allowPrivateAddresses, allowWideWildcards and wideWildcards
3 policyFile — a repo-local YAML file tighten only

Rank 3 is attacker-controlled: a hostile repository ships one, and a prompt-injected agent can write one. It may add deny patterns and raise audit to enforce. That is all:

v: 1
addDeny: ['*.internal.example', 'paste.example']
enforce: true

There is no allow, no way to open an address range, no way to name a wildcard namespace or to choose the wide-wildcard posture in either direction, no way to name the spool, and enforce: false is an error rather than an ignored key. Any other key, and any downgrade, makes the whole file invalid: it is reported on process.stderr and the deployment’s logger, then ignored, never obeyed in part. A missing file is not an error — the recommended policyFile is workspace-relative, so failing the mount would stop dsh from starting in every repository without one, and would let a hostile repository remove the control by shipping a broken file.

The file is parsed with js-yaml under JSON_SCHEMA, so a !!js/function tag is a parse error rather than code execution, and it never goes near the Cordis loader.

The harness takes the same line: packages/boot/app-boot/src/index.ts:111 forbids a repo-local .env from setting HTTP_PROXY / HTTPS_PROXY / ALL_PROXY / NO_PROXY.