Operations & production hardening
The bundled docker compose / Helm defaults are tuned for a demo: local database,
no auth, no TLS, best-effort persistence. This runbook is what changes for a real
deployment. It complements, and does not repeat, the threat model
(trust boundaries + operator checklist), the README “Application hardening” section, and
SECURITY.md.
Environment-variable names below are the actual configuration keys (see
backend/internal/config).
1. Demo vs production - what actually changes
Section titled “1. Demo vs production - what actually changes”| Concern | Demo default | Production |
|---|---|---|
| Database | bundled Postgres+AGE image | external, encrypted Postgres+AGE - managed only on Azure, self-managed elsewhere (§3) |
| API auth | open | API_TOKENS and/or OIDC (OIDC_*) required |
| Ingest auth | open | per-tenant HMAC (INGEST_HMAC_SECRETS) + INGEST_RATE_RPS |
| Transport | plaintext | TLS in-app (TLS_CERT_FILE/TLS_KEY_FILE) or at the ingress; POSTGRES_SSLMODE; NATS_TLS_* |
| Secrets | env vars in compose | secret manager / mounted files, never inline |
| Exposure | localhost | behind a gateway/WAF; network policy between components |
2. Secure configuration reference
Section titled “2. Secure configuration reference”Turn these on for any deployment reachable beyond a trusted boundary:
# API authentication (pick tokens, OIDC, or both)API_TOKENS=<token>:<role>[:<tenant>[:<YYYY-MM-DD>]],... # role is viewer|admin # e.g. API_TOKENS=s3cr3t:admin # Entries that do not parse are DROPPED, # so with PG_ENV=production the backend # refuses to start rather than serve an # open endpoint. Check startup warnings.OIDC_ISSUER=https://idp.example.com/realms/pg # + OIDC_CLIENT_ID, OIDC_JWKS_URL, # OIDC_AUDIENCE, OIDC_AUTHORIZE_URL, # OIDC_TOKEN_URL, OIDC_SCOPESAUTH_LOCKOUT_THRESHOLD=5 # brute-force lockout
# Ingest authentication + throttlingINGEST_HMAC_SECRETS=<tenant>:<hmac-secret>,... # per-tenant webhook HMACINGEST_RATE_RPS=50 # requests/sec cap
# Transport securityTLS_CERT_FILE=/etc/pg/tls/tls.crt # in-app TLS (or terminate at the ingress)TLS_KEY_FILE=/etc/pg/tls/tls.keyPOSTGRES_SSLMODE=verify-full # never `disable` in prodNATS_TLS_CA=/etc/pg/nats/ca.crt # + NATS_TLS_CERT, NATS_TLS_KEY
# Data hygieneSCRUB_INGEST=true # scrub secrets out of ingested payloadsCORS_ALLOWED_ORIGINS=https://dashboard.example.comLeast privilege for the outbound integrations: give connectors a read-only cloud role
(AWS_CONNECTOR_MODE=sdk + a SecurityAudit/ViewOnlyAccess role), scope GITHUB_TOKEN
to the target repo only, and leave ANTHROPIC_API_KEY/HF_TOKEN unset unless you accept
sending attack-path context to that provider.
Token scope is the second line, not the first. REPO_ALLOWLIST is the first, because
the destination of a forge write comes from an ingested node property rather than from
your configuration - so it is chosen by whoever can post an event, which includes every
scanner holding the ingest HMAC key. Name the repositories explicitly:
REPO_ALLOWLIST=acme/payments-api,acme/* # exact slugs, or an owner wildcardEmpty refuses every real write. The pairing matters: the allowlist bounds where the engine asks to write, the token scope bounds what the forge lets it write, and neither alone is sufficient.
Two ready-to-use hardened profiles apply all of the above:
-
Kubernetes (recommended):
deploy/helm/perspectivegraph/values-production.yaml- auth + ingest HMAC on, external Postgres+AGE withsslMode: verify-full, TLS at the ingress, durable audit log, and secrets sourced from your own manager (secrets.existingSecret, e.g. Vault / Sealed Secrets / External Secrets).Terminal window helm upgrade --install perspectivegraph oci://ghcr.io/luiacuaniello/charts/perspectivegraph \-f values-production.yaml \--set postgres.externalHost=db.internal --set ingress.host=pg.example.comThe chart installs from the registry, so a production cluster needs no git checkout. The profiles ship inside the chart, so neither does getting them - pull it once and the values files are there to edit and keep under your own change control:
Terminal window helm pull oci://ghcr.io/luiacuaniello/charts/perspectivegraph --untarls perspectivegraph/values-*.yaml # production, ha, sso-demoIt is cosign-signed; verify it before it runs (see the manual). From a git checkout instead, swap the
oci://…reference fordeploy/helm/perspectivegraph. -
Docker Compose (single host / on-prem):
.env.production.example(copy to.env, fill in,chmod 600) plus thedocker-compose.prod.ymloverride for the in-app TLS cert mount.Terminal window cp .env.production.example .env # then fill it in and: chmod 600 .envdocker compose --profile app -f docker-compose.yml -f docker-compose.prod.yml up -d
Keeping secrets out of the environment
Section titled “Keeping secrets out of the environment”Every credential above also accepts a <KEY>_FILE variant holding the path to a
file with the value: API_TOKENS_FILE, INGEST_HMAC_SECRET_FILE,
POSTGRES_PASSWORD_FILE, STORE_ENCRYPTION_KEY_FILE, EXPORT_SIGNING_KEY_FILE,
GITHUB_TOKEN_FILE, GITLAB_TOKEN_FILE, ANTHROPIC_API_KEY_FILE, HF_TOKEN_FILE, and
the rest.
Use them. The process environment is one of the least private places on a Unix host:
anything that can read /proc/<pid>/environ sees it, docker inspect prints it in full
to anyone in the docker group, it lands in crash dumps, and every child process
inherits it. A mounted file keeps the value out of all of that, and it is what makes
Docker secrets, Swarm secrets, a Vault Agent sidecar and a Kubernetes Secret mounted as a
volume work without the credential ever passing through the environment.
Docker. An overlay ships with the repo:
mkdir -p secrets && chmod 700 secretsprintf '%s' "$(openssl rand -hex 32)" > secrets/ingest_hmac_secretprintf '%s' "tok-$(openssl rand -hex 16):admin" > secrets/api_tokensprintf '%s' "$(openssl rand -hex 32)" > secrets/postgres_passwordprintf '%s' "$(openssl rand -hex 32)" > secrets/store_encryption_keychmod 600 secrets/*
docker compose -f docker-compose.yml -f docker-compose.secrets.yml --profile app up -dIt mounts each file at /run/secrets/…, points the matching *_FILE variable at it,
blanks the environment-borne version so a stale value cannot win, and sets
PG_ENV=production. secrets/ is gitignored. Verify with
docker inspect perspective-backend | grep -i secret - you should see only paths.
Kubernetes. The chart already provisions a Secret and consumes it with
secretKeyRef; set secrets.existingSecret to bring your own from External Secrets,
Sealed Secrets or Vault. To go further and avoid the environment entirely, mount that
Secret as a volume and set the *_FILE variables to the mounted paths.
A <KEY>_FILE that is set but unreadable stops the process. That is deliberate: an
operator who mistypes a mount path would otherwise start cleanly with no credential at
all - and “no API token” is not an error, it is the demo profile - which is exactly the
open deployment the mount was meant to prevent. A trailing newline is stripped (secret
managers add one); nothing else is touched, so a passphrase may begin or end with a space.
3. The database: PostgreSQL + Apache AGE
Section titled “3. The database: PostgreSQL + Apache AGE”The graph lives in PostgreSQL with the Apache AGE extension, and that extension - not the DSN - is what decides your deployment. Most managed PostgreSQL services do not offer AGE, so “point it at a managed Postgres+AGE” is not advice you can act on everywhere, and it is worth settling before anything else in this runbook.
Where you can actually get one
Section titled “Where you can actually get one”| Where | AGE | What it costs you |
|---|---|---|
| Azure Database for PostgreSQL flexible server | yes (PostgreSQL 16 and below) | The only managed service that ships it. Turned on with two server parameters (recipe below). Not available on PostgreSQL 17, and AGE is excluded from Azure’s in-place major-version upgrade |
| Self-managed on Kubernetes or a VM | yes | Backups, failover, patching and TLS become yours. The apache/age image, or the one this project builds in deploy/postgres, already preloads the extension |
| AWS RDS / Aurora PostgreSQL | no | Not on the extension allow-list. Requests go to rds-postgres-extensions-request@amazon.com |
| Google Cloud SQL / AlloyDB | no | Not in the supported-extensions list |
The bundled perspectivegraph-postgres container |
yes | Demo only. Not sized, backed up or tuned for production, whatever its vulnerability report says |
So on AWS and GCP today the honest choice is to run Postgres+AGE yourself. One thing makes that a smaller decision than it looks: the graph is derived state - every node and edge is reconstructible by re-ingesting the feeds - so a lost database is a re-seed rather than a data-loss event (§4). Losing it costs you history, not the map.
What the engine needs of it
Section titled “What the engine needs of it”Tested against PostgreSQL 16 + AGE 1.6.0 - the digest-pinned image the demo and the CI integration job both run. Other combinations are not tested here.
The role does not need to be a superuser. These grants are enough:
GRANT CONNECT, CREATE ON DATABASE perspectivegraph TO perspective;GRANT USAGE ON SCHEMA ag_catalog TO perspective;GRANT SELECT ON ag_catalog.ag_graph, ag_catalog.ag_label TO perspective;What the connection does need is AGE actually loaded, and there are only two ways that happens:
- the role may run
LOAD 'age'- a superuser-only command, which is what the bundled demo does and what no managed service permits; or ageis inshared_preload_libraries, so the library is present before the session opens. The bundled image does this by default; Azure exposes it as a parameter.
The backend works out which of the two it is on the first query and adapts. Check what you have with:
SELECT extversion FROM pg_extension WHERE extname = 'age'; -- is it installedSHOW shared_preload_libraries; -- is it preloadedLet the backend create its own graph. AGE keeps each graph in a schema owned by
whoever called create_graph, and a role that does not own that schema is refused on the
first write with permission denied for schema .... If an administrator pre-creates the
graph, hand it over with ALTER SCHEMA <graph> OWNER TO <role> (and the tables in it)
rather than leaving the engine a tenant in someone else’s schema.
Azure Database for PostgreSQL
Section titled “Azure Database for PostgreSQL”- On the server’s Parameters blade, add
AGEtoazure.extensionsand toshared_preload_libraries, and save. The server restarts to load the library. - Connect to the database and run
CREATE EXTENSION IF NOT EXISTS age CASCADE;. - Apply the grants above and point the backend at it.
Do not run LOAD 'age' by hand there: with the library preloaded it does not quietly
succeed, it fails with a privilege error.
Self-managed
Section titled “Self-managed”Run ghcr.io/luiacuaniello/perspectivegraph-postgres (or apache/age, or your own build
of the extension) as a StatefulSet, under a
PostgreSQL operator with that image, or on a VM. Whatever you pick, the list of things you
have just taken on is the same, and none of it is optional for production: backups with a
tested restore (§4), a replica and a failover path, patching for both PostgreSQL and AGE,
TLS, and monitoring. §4 covers backup and restore; the demo image is not a starting point
for any of it.
Pointing the backend at it
Section titled “Pointing the backend at it”POSTGRES_DSN=postgres://user:pass@db.internal:5432/perspectivegraph?sslmode=verify-full# or the discrete POSTGRES_HOST/PORT/DB/USER/PASSWORD + POSTGRES_SSLMODE keys# POSTGRES_PASSWORD_FILE keeps the password out of the environment entirely (§2)4. Backup & restore (the graph is sensitive data)
Section titled “4. Backup & restore (the graph is sensitive data)”The graph in Postgres+AGE is your source of truth and a map of the attack surface - back it up and test the restore.
# Backup: dump the whole database (includes ag_catalog + the graph schema)pg_dump --format=custom --no-owner --dbname="$POSTGRES_DSN" --file pg-graph.dump
# Restore into a fresh instance that already has the AGE extension loadedcreatedb perspectivegraphpsql -d perspectivegraph -c 'CREATE EXTENSION IF NOT EXISTS age;'pg_restore --no-owner --dbname=perspectivegraph pg-graph.dumpNotes:
- Restore into a database where
CREATE EXTENSION agehas run first; AGE graph data lives underag_catalogand the graph’s own schema, both captured by a fullpg_dump. - Store dumps encrypted (they contain A1 from the threat model). Apply the same retention and access controls you would to a secrets store.
- Validation/verdict data persists separately via
VALIDATIONS_PATH; back that path up too if you rely on the calibration history.
5. Upgrades
Section titled “5. Upgrades”- Read UPGRADING.md for the target version - it lists the releases that need an action from you, and what the action is. Then the CHANGELOG for everything else that changed.
- Take a backup (section 4).
- Roll the backend image forward. The graph schema is created/managed by the backend; there is no separate migration step, but a major version may re-derive nodes/edges - a backup lets you roll back.
- Verify:
/healthzreturns 200 andattackPathsreturns after oneANALYZER_INTERVAL.
Pin to a signed, digest-referenced image and verify it before rollout (see SECURITY.md “Our own supply chain”).
6. Observability & SLOs
Section titled “6. Observability & SLOs”GET /healthz- liveness/readiness (the container HEALTHCHECK uses thehealthzsubcommand; distroless has no shell).GET /metrics- Prometheus metrics:perspectivegraph_connector_*,perspectivegraph_analyzer_*, ingest and auth counters, andperspectivegraph_broker_connected(0 while the bus is reconnecting).perspectivegraph_auth_denied_total{surface,reason}counts refused credentials where they are refused (API and ingest;anonymous_aiandanonymous_gateare visitors turned away from the paid and the heavy endpoints);perspectivegraph_ingest_signatures_total{version}shows whether anything still signs v1;perspectivegraph_graph_pending_edgescounts edges waiting for an endpoint. The pending count is refreshed by each replica’s own analyzer pass, so on a replica that wrote nothing it can lag by a few passes.perspectivegraph_graph_sweeps_total{source,result}counts complete snapshots applied andperspectivegraph_graph_swept_total{kind}what they removed (node,edge) or set waiting (parked); a jump in removals for one source is the thing to look at when an asset disappears unexpectedly - the log line names the scope.- Suggested SLOs to alert on: ingest error rate, analyzer pass duration vs
ANALYZER_INTERVAL, connectorlast_error, andauth.denyspikes (possible credential stuffing) from the audit log. - Ready-to-use:
deploy/observabilityships a Grafana dashboard and Prometheus alert rules for exactly these signals - import the dashboard, load the rules, point a scrape at/metrics. For analyzer load/scale characterization see SCALE.md (make scale-test).
Health, and what it now refuses to hide
Section titled “Health, and what it now refuses to hide”GET /healthz returns 503 when the engine is serving in a reduced mode, and 200
otherwise. The Helm readiness probe hits it, so a degraded pod is taken out of rotation.
That distinction did not exist until it was found by running the stack: /healthz
returned 200 unconditionally, so the probe only ever proved the HTTP server was
listening. Meanwhile, when Apache AGE was unreachable at startup the backend fell back
to in-memory stores with a single warning - leaving the engine computing over an empty,
volatile graph. An empty graph answers “no attack paths”, which reads as good news. So
a database outage presented as a clean bill of health, one layer below where
ingestCoverage can see it.
Two changes close that:
PG_ENV=productionnow refuses to start when Apache AGE is unavailable, instead of falling back. SetPG_ENV=demoif you genuinely want the in-memory store, orGRAPH_STRICT=trueto get the same refusal outside production.- When the fallback does engage (demo profiles),
/healthzreports 503 with the reason, so nothing downstream mistakes it for a working deployment. A deployment that runs in-memory by design stays healthy - that is the demo working as intended, not a failure.
/healthz does not fail while the event bus is down, on purpose: the API still
serves the graph it has, and taking every replica out of rotation for a NATS restart
would turn a pause in ingest into an outage of the dashboard. The bus is covered in two
other ways instead:
- A reconnect is retried for as long as it takes, and on reconnecting the backend
recreates its stream and consumer if the server came back empty. Watch
perspectivegraph_broker_connected; the shippedPerspectiveGraphBusDisconnectedalert fires when it stays 0. - What cannot recover ends the process. A listener that cannot bind (a taken port, an unreadable certificate), a connection closed for good, or a consumer that stops exits with status 1 and the reason on the last log line, so Kubernetes or Compose restarts it. It used to log once and keep running without that part of itself - and the liveness probe, which only checks that the API port accepts connections, saw nothing wrong.
- Set
LOG_FORMAT=jsonin production. The default istext, which is what a person wants duringmake demoand what no log pipeline wants.LOG_LEVELisdebug|info|warn|error. Both go to stdout; collect them there rather than writing files. - Every request carries an id. It is generated per request (or taken from an inbound
X-Request-Idwhen that value is short and alphanumeric), returned in theX-Request-Idresponse header, attached to log lines made with the request’s context, and written into the audit record’sfields.request_id. So one identifier joins what a user reports, what the application log says, and what the audit log recorded - which is the join you want at the moment you actually need it.- It rides in the audit record’s
fields, not in a new column, so the hash chain still verifies exactly as before andverify-auditis unaffected. - Alerts raised by the abuse watchers (
exfil.alert,auth.lockout.alert) deliberately carry no request id: they describe a window of events crossing a threshold, not the one request that happened to be last, and pinning them to it would point an investigation at an arbitrary call.
- It rides in the audit record’s
/metricsis open and unthrottled - deliberately, so a scrape never starves. Series such asperspectivegraph_analyzer_critical_pathscarry atenantlabel, so on a reachable port they are enough to enumerate tenants and read each one’s current path count.- Set
METRICS_ADDRto move them off the API port onto their own listener, e.g.METRICS_ADDR=127.0.0.1:9090. That listener serves/metricsand nothing else - no GraphQL, no auth config, no exports - and speaks plain HTTP by design, because it is meant for an address the outside cannot reach and demanding a certificate there is the friction that pushes operators back onto the public port.values-production.yamlsets it. - It is empty by default:
/metricson the API port is declared stable surface in API-STABILITY.md, and relocating it silently would break every existing scrape config.
- Set
7. High availability
Section titled “7. High availability”The analyzer/scheduler and connectors are leader-gated - extra replicas do not duplicate work or multiply API calls - and leadership fails over on its own. The election is a session-scoped PostgreSQL advisory lock: when the holder dies its connection drops, the lock is released server-side, and another replica takes it on its next check. No external coordinator, and no operator action.
What pins you to one replica is the file governance backend, not the election. With
GOVERNANCE_BACKEND=postgres the suppressions, tickets, posture history, validations and
KEV holdout are all shared - and so is the audit log: the chain moves into the
database, where each append takes a transaction-scoped advisory lock, reads the tail and
writes in the same transaction, so several replicas extend one chain instead of forking
it. Keep the database HA at the managed-Postgres layer.
So the Kubernetes recipe for more than one backend is governanceBackend: postgres and
persistence.enabled: false - the PVC then holds nothing, and its ReadWriteOnce mode would
otherwise leave every pod scheduled off the first node hanging on FailedAttachVolume. The
chart refuses to render persistence.enabled with replicas > 1 rather than let you find
that out from a stuck rollout.
values-ha.yaml is that recipe as an overlay on top of the production profile, so HA is
four decisions rather than a second copy of the whole file:
helm upgrade --install perspectivegraph oci://ghcr.io/luiacuaniello/charts/perspectivegraph \ -f values-production.yaml -f values-ha.yaml \ --set postgres.externalHost=db.internal --set ingress.host=pg.example.com \ --set nats.externalUrl=nats://nats.internal:4222It sets the two settings above, raises the backend to three replicas and the dashboard to
two, spreads them one-per-node, and adds a PodDisruptionBudget - because replicas: 3
that a drain can evict all at once, or that the scheduler stacks on one node, is three
copies of a single failure domain rather than availability. values-production.yaml keeps
the single-replica shape on the file backend: fewer moving parts, and the audit chain in a
file you can archive.
Two single points of failure it does not remove, and neither is hidden:
- The database. Every replica shares one Postgres+AGE. Use a managed instance with a replica and automatic failover - §3 covers where AGE is actually available, which is narrower than it looks.
- The event bus. The bundled NATS is one replica whose JetStream store lives on an
emptyDir, so rescheduling the pod drops in-flight events (the backend recreates the stream on reconnecting, so ingest resumes on its own). The HA overlay therefore refuses to inherit it: it setsnats.enabled: falseand the chart will not render untilnats.externalUrlpoints at a NATS you run clustered - and give the stream replicas there (nats stream edit PERSPECTIVE --replicas 3): the backend sets only the stream’s subjects, retention and age limit, and keeps every other setting you make. What that outage costs is the events in flight, not the graph - the graph is derived and the feeds re-ingest it - but the analyzer is blind for the duration.
Retention and rotation
Section titled “Retention and rotation”An append-only chain nobody may delete from grows forever, and “forever” is the absence of a retention policy rather than one. Both backends can be bounded, differently:
Postgres chain - AUDIT_RETENTION. Set a window and the leader prunes records older
than it, oldest first, on a cadence derived from the window (a sixth of it, clamped to
between an hour and a day - so a 90-day window checks in daily):
AUDIT_RETENTION=2160h # 90 days; unset (the default) keeps everything, as beforePruning removes a prefix and records a checkpoint holding the sequence and hash of the
last record removed, so what survives still verifies link-by-link back to it - see the
threat model for why that is the only
shape of deletion this log allows. verify-audit says what was pruned rather than quietly
reporting a shorter chain, the prune is itself an entry in the chain it shortened, and
perspectivegraph_audit_pruned_records_total counts what has gone.
File chain - rotate it. There is no automatic pruning, and the two obvious ways to
rotate are both wrong: logrotate with copytruncate leaves the process writing at its old
offset into a truncated file, and a live mv leaves it writing into the file you meant to
retire, because the open handle follows the inode. The procedure that works is stop, move
the file aside, start - the engine then begins a fresh chain, and each retired file stays
a complete chain that verifies on its own (a test pins this). Archive the retired files
somewhere append-only: once rotated, their integrity is your archive’s problem, not the
engine’s. Setting AUDIT_RETENTION on a file-backed deployment logs a warning and prunes
nothing, rather than pretending to.
Whichever you use, decide the window on purpose. It is the number a data-protection review asks for, and the erasure/tamper-evidence tension in the threat model is the reason it cannot simply be “delete that one record”.
Verifying the chain
Section titled “Verifying the chain”Verify the chain wherever it lives - the subcommand follows it:
perspectivegraph verify-audit /var/log/perspectivegraph/audit.log # file backendperspectivegraph verify-audit -postgres # governance database-postgres reads POSTGRES_DSN (or the discrete POSTGRES_* keys, _FILE variants
included) exactly as the server does, so a DSN with a password in it never has to be typed
into a command line where every process on the host can read it.
8. Air-gapped and internal-registry installs
Section titled “8. Air-gapped and internal-registry installs”Nothing here phones home: the engine opens no outbound connection until you set a key or a flag, so an isolated network is a supported deployment rather than a workaround. What an air-gapped install needs is the artefacts brought in, and their signatures checked before they cross the boundary - verifying inside is verifying a copy you already trusted.
Mirror, verifying on the way in. On a host with internet access:
# The release you are approving.V=1.26.0 # x-release-please-versionID_RE='https://github.com/luiacuaniello/perspectivegraph/.*'ISSUER=https://token.actions.githubusercontent.com
for img in perspectivegraph perspectivegraph-dashboard; do cosign verify --certificate-identity-regexp "$ID_RE" --certificate-oidc-issuer "$ISSUER" \ "ghcr.io/luiacuaniello/${img}:v${V}" # Copy by DIGEST so the tag cannot be re-pointed between verification and pull. digest="$(crane digest "ghcr.io/luiacuaniello/${img}:v${V}")" crane copy "ghcr.io/luiacuaniello/${img}@${digest}" "registry.internal/perspectivegraph/${img}:v${V}"done
cosign verify --certificate-identity-regexp "$ID_RE" --certificate-oidc-issuer "$ISSUER" \ "ghcr.io/luiacuaniello/charts/perspectivegraph:${V}"helm pull "oci://ghcr.io/luiacuaniello/charts/perspectivegraph" --version "$V"Then point the chart at your registry:
helm install perspectivegraph "./perspectivegraph-${V}.tgz" \ -f perspectivegraph/values-production.yaml \ --set backend.image.repository=registry.internal/perspectivegraph/perspectivegraph \ --set frontend.image.repository=registry.internal/perspectivegraph/perspectivegraph-dashboardThree things that would otherwise try to leave, all off by default - check they still are, because an air-gapped cluster turns a silent outbound call into a hang rather than an error:
| Setting | Off by default | What it would reach |
|---|---|---|
THREATINTEL |
yes (off) |
CISA KEV + FIRST EPSS feeds |
ANTHROPIC_API_KEY / HF_TOKEN |
yes (unset) | the model provider - and it sends attack-path context |
GITHUB_TOKEN |
yes (unset) | api.github.com, for PR comments and the merge gate - bounded by REPO_ALLOWLIST |
The two images and the chart are the whole dependency set. The engine ingests what you POST to it, so no scanner needs outbound access either - only a route to the ingest port.
The database is the exception worth planning for. Apache AGE has to come from somewhere: mirror the bundled database image, or have your DBA team build the extension for your managed instance (§3).
9. Continuous delivery (Argo CD, Flux)
Section titled “9. Continuous delivery (Argo CD, Flux)”The chart is an OCI artefact, so a GitOps tool consumes it directly - no repository to clone, no chart to vendor:
apiVersion: argoproj.io/v1alpha1kind: Applicationmetadata: name: perspectivegraph namespace: argocdspec: project: default source: repoURL: ghcr.io/luiacuaniello/charts chart: perspectivegraph # Pinned: let a bump be a reviewed commit, not a surprise resync. targetRevision: 1.26.0 # x-release-please-version helm: valueFiles: [values-production.yaml] # ships inside the chart; override with your own destination: server: https://kubernetes.default.svc namespace: perspectivegraph syncPolicy: automated: { prune: true, selfHeal: true }Keep the secrets out of it: secrets.existingSecret points at a Secret your own operator
(External Secrets, Sealed Secrets, Vault) manages, so nothing sensitive lives in the
Application manifest.
There is no Terraform module. The chart plus a helm_release resource is the whole
integration, and a module wrapping that would be a layer to maintain rather than a
capability to gain.
10. Pre-production checklist
Section titled “10. Pre-production checklist”- API auth enabled (
API_TOKENS/OIDC) and verified from an unauthenticated client. -
GET /auth/mewith each issued token answers the role you meant to grant. - Ingest HMAC (
INGEST_HMAC_SECRETS) +INGEST_RATE_RPSset, and - onceperspectivegraph_ingest_signatures_total{version="v1"}stays at 0 -INGEST_HMAC_ACCEPT_V1=false, so a captured ingest request cannot be replayed. - TLS everywhere (
TLS_*,POSTGRES_SSLMODE=verify-full,NATS_TLS_*). - External Postgres+AGE chosen with §3 open (managed on Azure, self-managed on AWS/GCP), its role non-superuser, and the demo image out of the deployment.
- Secrets in a manager/mounted files, not inline in compose/Helm values.
- Connector role is read-only and reviewed;
GITHUB_TOKENscoped to one repo, andREPO_ALLOWLISTnames the repositories it may write to (empty = no writes). - Backup scheduled and a restore rehearsed (section 4).
-
/metricsscraped; alerts on the SLOs above. - Images verified (cosign signature + SBOM + provenance) before rollout.
- Engine behind a gateway/WAF; network policy between components (
networkPolicy.enabled, andnetworkPolicy.backendFromnaming the ingress controller’s namespace). - NATS authenticates its clients: the chart’s does on its own; an external one needs
NATS_USER/NATS_PASSWORD. Rendering withhelm template(Argo CD, Flux), pinpostgres.auth.passwordandnats.auth.password(or bring the Secrets), since a render without the cluster cannot keep a generated one.
See the threat model operator assumptions for the rationale behind each item.
11. Publishing a read-only instance
Section titled “11. Publishing a read-only instance”The opposite deployment to §10: one meant to be reached by people who have no credential - a public demo, or an internal dashboard a whole company may read. The engine supports it directly rather than leaving it to a proxy rule, because a rule that lives outside the binary is one no test holds.
# API_ANONYMOUS_ROLE=viewer: a caller with NO credential gets the viewer role.docker compose -f docker-compose.yml -f docker-compose.demo.yml -f docker-compose.public.yml \ --profile app up -dOn Kubernetes the same switch is auth.anonymousRole: viewer.
What the backend enforces. viewer is the only value it accepts: anything that can
write must be tied to a credential, and a typo (admin, viewr) stops the process at
startup rather than being quietly downgraded. Writes stay admin-only, so an anonymous
suppression, verdict, ticket or remediation PR answers 403. A presented-but-wrong token
still fails with 401 - it does not fall through to anonymous, or revoking a leaked token
would silently demote its holder to public read instead of locking them out.
What the dashboard shows. GET /auth/config answers authRequired: false with
anonymousRole: "viewer", so the dashboard opens without a sign-in and carries a quiet
read-only instance notice, which is also what tells a visitor why a suppression is refused.
An instance left open by accident reports no anonymousRole and keeps the red
open-instance banner. The backend tells the two apart; the page does not guess.
What it does not do.
- It does not make the data safe to publish. Everything the dashboard shows - assets,
versions, reachable routes to crown jewels - becomes public. An attack map is precisely
what an attacker would ask for. Seed a published instance with sample data
(
make seed), never with a real estate’s findings. - It does not close the write side.
/ingestis how the graph is built; never proxy port 8081. Everything indocker-compose.ymlalready binds to 127.0.0.1, so publish only the dashboard (3000) through your TLS terminator and nothing else. - It does not slow anything down for you. The costly queries (
whatIfre-runs the simulation,kShortestPathsenumerates routes) are now reachable without a credential, so the override lowersAPI_RATE_RPSto 10 per client IP. Each request may run at most 20 of those analyses and half the cores run them at once, so one visitor cannot take the machine;perspectivegraph_api_heavy_refused_totalcounts what was turned away. - It does not answer AI questions for visitors.
/ai/*needs a signed-in caller, so anonymous visitors get none even if a key is configured - but keep keys off a published instance anyway. - It does not run the merge gate for visitors.
POST /gate/impactanalyses a report the caller sends, so it needs a signed-in caller too; a CI gate pointed at a published instance passes a token.
Per visitor, not per proxy. A published instance is always reached through a proxy:
in the compose recipe every visitor comes through the dashboard’s nginx. Keyed on that
peer, the rate limit and the brute-force lockout are one key for everybody, and fifty wrong
tokens from one person lock every visitor out. That was reproduced on a real stack. So the
override sets TRUSTED_PROXY_CIDRS=172.16.0.0/12, the range Docker gives docker0 and the
compose networks after it, and the backend reads the real client from X-Forwarded-For.
It believes only the hops those addresses appended, so a visitor cannot choose a key or
aim a lockout at someone else. Two conditions:
- Your TLS proxy must set
X-Forwarded-For. Caddy, nginx and Traefik do by default; a plain TCP forwarder does not. - The compose network must be in that range. Check with
docker network inspect <project>_default; if your daemon allocates elsewhere (a custom address pool, or more than fifteen networks on the host), setTRUSTED_PROXY_CIDRSto that subnet.
On Kubernetes the ingress controller is the proxy: set backend.trustedProxyCidrs to its
pod CIDR.
Checklist for a published instance
-
API_ANONYMOUS_ROLE=viewer, and an unauthenticatedPOST /suppressionsanswers 403. -
/auth/configshows"anonymousRole":"viewer": the dashboard shows the read-only notice, not the red open-instance banner. - Sample data only; no connector credentials, no
GITHUB_TOKEN, no AI keys. - Only the dashboard port is proxied;
/ingestunreachable from the internet. - TLS at the proxy, and
API_RATE_RPSlow. - The backend logs
trusted proxies configured, and the proxy CIDR covers the network the dashboard runs on. - An
API_TOKENSadmin credential kept for yourself, if you need to change anything;/auth/mewith it answers"canWrite": true.
An MCP client can be pointed at a published instance the same way the dashboard is: the
server is a client of this API, so perspectivegraph mcp --api https://<host> answers from
it without any credential.