Skip to content

Upgrade notes

What changes for a running deployment when you move between versions - the settings you have to set, the requests that start failing, the behaviour that is no longer what it was.

This file exists because the CHANGELOG does not carry it. That file is generated from Conventional Commit subjects, so it names what changed in one line and drops the paragraph underneath explaining what an operator has to do about it. A one-line entry is fine for a bug fix and useless for a release that refuses writes until a new variable is set.

Only versions that need an action appear here. A version missing from this list is one you can upgrade into without touching your configuration - which is most of them. Follow the upgrade recipe in SUPPORT.md either way: pin by digest, take the backup, stage it.


P1 means “act now”, and red-team verdicts move a route

Section titled “P1 means “act now”, and red-team verdicts move a route”

Affects you if you triage by band (P1/P2/P3), read priority or priorityLabel from the API or the MCP tools, or alert on P1 counts.

A route with a runtime alert into AdministratorAccess used to read P2: the blended priority weighs exploitability at a third and could not reach P1 without a KEV entry. Now:

  • A fact about the route makes it P1: a runtime alert on it, an asset open to anyone, a KEV weakness on a route an attacker is likely to complete, or a likely route (≥ 80%, on evidence rather than default weights) into a high-value asset. priorityReason (new) says which. Every P1 path sits at 70 plus 30% of its blended priority, so the band keeps an order and a natural P1 of 76 now reads 92.8.
  • Red-team and BAS verdicts re-band a route where paths are listed (attackPaths, the dashboard, the MCP tools): confirmed → P1; refuted → P3, unless a runtime alert or open access contradicts the test - then it stays and says the evidence conflicts. The Priority recorded with a verdict, which grades the triage order, still contains none.
  • Expect more P1s where routes are runtime-confirmed or tested; expect refuted routes at the bottom. The dashboard no longer re-sorts refuted routes itself.

AFFECTS (a library has a CVE) and DEPENDS_ON (an image ships a library) no longer carry a technique: they are facts about software, not attacker actions. An exploit is T1190 (initial access) until the attacker is inside, then T1210 (lateral movement), and the exposure before an exploit carries no technique of its own. CAN_ESCALATE_TO is T1098.003 (Account Manipulation: Additional Cloud Roles), no longer T1078.004. Anything that stored or keyed on the old technique ids per hop will see different ones.

Calibration reads that could not be trusted

Section titled “Calibration reads that could not be trusted”

brier_recalibrated is measured out of sample at every size - leave-one-out below 20 outcomes, where it was the in-sample fit and read a perfect 0.000 on a dozen points. Below 30 outcomes the Trust page headlines “Not enough outcomes yet” with the direction beside it, and an AUC below its floor shows the counts it rests on instead of an interval.

Action: none; nothing to configure. Check anything that alerts on P1 counts or keys on per-hop ATT&CK ids.

The chart’s bus authenticates and keeps its stream, and its database password is no longer shared

Section titled “The chart’s bus authenticates and keeps its stream, and its database password is no longer shared”

Affects you if you install with the Helm chart.

The chart’s NATS accepted any client, so any pod that could reach it published events straight into the graph - past the ingest webhook’s signature check - or deleted the stream. It kept that stream on an emptyDir, so a restart lost every queued event. And the bundled Postgres used the password perspective on every install.

What changes on helm upgrade:

  • NATS requires a user and password. The chart generates the password into a Secret of its own (<release>-perspectivegraph-nats-auth), gives it to the backend, and keeps it across upgrades. NATS and the backend both roll during the upgrade, and ingest pauses until both are up. An external NATS is untouched unless you set nats.auth.user and nats.auth.password (or nats.auth.existingSecret); the backend reads NATS_USER/NATS_PASSWORD in any deployment.
  • NATS is a StatefulSet with a 2 GiB volume (nats.persistence). The upgrade replaces the Deployment, so events still queued in the old pod are lost this once - as every NATS restart lost them before. The volume needs a default StorageClass, as the bundled Postgres’s already does; nats.persistence.enabled: false keeps the emptyDir.
  • postgres.auth.password defaults to empty. A new install gets a random password. An existing one keeps the password in its Secret - perspective unless you set one - because Postgres reads it only when it first creates its data directory. To rotate it, ALTER ROLE in the database, then set postgres.auth.password.
  • secrets.existingSecret works for the bundled Postgres. Its pod named the chart’s own Secret, which does not exist when you bring yours, so it could not start. It now reads yours, as the backend always did.
  • networkPolicy: only the backend may reach the bundled NATS and Postgres. Off by default, on in values-production.yaml; backendFrom also closes the backend to all but the dashboard and the peers you name.

Action if you render with helm template (Argo CD, Flux): a render without the cluster cannot read the existing Secrets, so it would draw new passwords on every sync - and a new Postgres password locks the backend out of the database it already has. Before upgrading, set postgres.auth.password to the current one (perspective if you never set it) and nats.auth.password, or bring both through secrets.existingSecret and nats.auth.existingSecret. An external Postgres now needs its password set explicitly there too.

Action to keep the old behaviour (not recommended): nats.auth.enabled: false, nats.persistence.enabled: false.

The graph forgets what its sources stop reporting

Section titled “The graph forgets what its sources stop reporting”

Affects you if you ingest Trivy or Semgrep reports, run the AWS connector, or rely on something staying in the graph after its source stopped listing it.

Until now the only way out of the graph was GRAPH_TTL, off by default, so a CVE fixed in an image, an instance terminated or a role deleted stayed - with the attack paths through it - until someone pruned. The engine now records who asserted each node and edge. When a source that describes a scope in full sends that scope again, whatever it said before and no longer says is withdrawn, once the whole ingest has landed. An element leaves the graph when no source asserts it any more.

What changes:

  • Trivy scans of an image by reference, Semgrep reports with ?repo=, and the AWS connector (per account; the network per account and region) are complete snapshots without any change on your side. A fixed CVE disappears at the next scan.
  • Everything else is unchanged unless you say so with ?snapshot=<scope> on the ingest webhook - ?snapshot=cluster:prod on a dump of the whole cluster. Name only what you really send in full: what the scope omits is removed.
  • A Trivy scan filtered by severity or fixability is complete for its filter only. If one pipeline sends the same image filtered and another unfiltered, they now undo each other’s findings: send the filtered one with ?snapshot=none.
  • A pull request’s scan never removes anything, and ?snapshot= together with slug/sha/pr is refused with 400.
  • Elements already in the graph when you upgrade are never removed this way - nothing recorded who sent them - only by GRAPH_TTL. To clean up leftovers from before, enable GRAPH_TTL for a cycle, or re-create the graph and let the feeds rebuild it.
  • The Postgres graph gains a provenance table beside the parked edges (<graph>._pg_provenance, plus _pg_removals). It is created and filled on first write, under the graph’s write lock, and needs no migration step. Every write now also records its origin, one row per element per source.
  • With ANALYZER_INCREMENTAL=true, the analyzer re-reads the whole graph after any removal, on every replica.
  • The Helm ingress now raises the nginx-based controllers’ 1 MiB request-body limit to the 32 MiB the backend accepts, for /ingest and /gate (ingress.maxBodySize). A body-size annotation you set yourself wins.

Action to keep the old behaviour: GRAPH_SWEEP=false (Helm graph.sweep: false): nothing is removed except by GRAPH_TTL.

The merge gate counts what the change adds, and writes nothing

Section titled “The merge gate counts what the change adds, and writes nothing”

Affects you if you run the merge gate - the GitHub Action, perspectivegraph gate, or the Trivy plugin.

The gate used to count every critical path through an asset stamped with the commit. A pull request that rescanned an image already in production, or re-rendered a deployment’s manifests, carried the commit onto assets that were on routes long before it - and every one of those routes blocked it: in the demo lab, nine paths where the change itself opened two. It now applies the report to a copy of the estate and counts only the routes the change opens or makes likelier. The routes it merely touches are reported (preexisting) and do not count. Nothing is written: the comparison runs on POST /gate/impact, on the API port, with the bearer token.

What changes for a running pipeline:

  • Fewer red checks, by design: the ones that remain are routes the change caused. max-critical now counts those.
  • Server mode no longer ingests the report. The live graph does not record a pull request’s scan unless you ask: persist: true (CLI -persist) sends it to the webhook after the verdict. The engine’s own pull-request comments and commit status are driven by what is ingested, so they need persist - or the webhook step you already had.
  • Server mode needs token for the comparison if API auth is on (the comparison is an authenticated API call, refused to anonymous callers on a public instance). ingest and hmac-secret are needed only with persist or attribution: commit.
  • Local mode compares with the estate you give it. Add base-reports - the scan of what runs now - or the routes through the scanned image count as the change’s, since the estate knows none of its findings.
  • A proxy in front of the API must pass /gate/ - the bundled dashboard nginx and the Helm ingress do - and let a report through: raise an ingress controller’s 1 MiB body limit (ingress-nginx: nginx.ingress.kubernetes.io/proxy-body-size: 32m), as /ingest already needed.

Action to keep the old behaviour: attribution: commit (CLI -attribution commit). The gate also falls back to it by itself - saying so in the log and in the attribution output - when there is no report to compare, or the engine predates 1.22. A gate binary older than 1.22 under the new action runs per commit with a warning. prVerdict is unchanged.

The engine’s commit status and PR comments count what the gate counts

Section titled “The engine’s commit status and PR comments count what the gate counts”

Affects you if the engine itself posts to your pull requests - GITHUB_TOKEN or GITLAB_TOKEN set, with REPO_ALLOWLIST.

The commit status perspectivegraph/attack-paths and the PR/MR comments counted every critical path through an asset stamped with the commit, like the gate before this release. They now apply the gate’s rule, comparing each analysis pass with the one before it. A route that is new, or likelier, belongs to the pull-request commits on it that arrived in between. A route that was already there when a commit arrived gets no comment and does not turn its status red, and the status description says how many there were. A route that appears later through a commit’s assets belongs to the pull request that arrived with it. If none did, it counts against the commits already on it, as “since this change arrived”.

What changes:

  • Fewer red statuses and fewer comments, by design, and the words change: “N critical attack path(s) opened or made likelier by this change”. A commit already in the graph when the engine starts - after a restart or an upgrade - has no “before”: every route through it counts, and the status says “in the graph before the engine was watching”. That is the gate’s rule for a commit the engine already holds.
  • Every commit on a route is judged, not only the first one found on it. Under the new rule the first could be innocent and the route belong to the second.
  • A status is posted when it changes, not re-posted on every analysis pass. GitHub keeps at most 1,000 statuses per commit and context.
  • Fixed: with more than one tenant, one tenant’s analysis pass posted success on another tenant’s red commits, and the last route of a tenant closing never cleared its red status.

Action to keep the old behaviour: PR_ATTRIBUTION=commit (Helm prAttribution: commit), which restores the old rule and wording. The three changes above stay.

Every probability the engine prints is recomputed under corrected rules, so the dashboard’s numbers move once on upgrade. Nothing needs configuring; what follows is what moves and why, so a step in a trend line is not mistaken for a change in the estate.

A fix is “verified” only when it protects something

Section titled “A fix is “verified” only when it protects something”

Affects you if you read the verified badge on remediations, verification over GraphQL, or whatIf results.

A fix’s verification compared two simulations of 800 trials that drew their random numbers independently, and called the fix verified when they differed by 0.05 points - against a noise of about ±2.5. A fix cutting an edge that leads nowhere was verified about half the time, and a what-if could show the risk rising after a cut. Both simulations now run the same trials, so the difference is the cut’s own effect: never negative, and exactly zero when the cut protects nothing.

Verification also reads a new measure, expectedReduction - the drop in the expected number of sensitive assets compromised - on verification and on whatIf. The old riskReductionPct / riskReduction (P(any asset compromised)) stays, but it saturates: while one asset is compromised in every trial it reads 0 for every other fix.

Action: none. Expect some fixes to turn from verified to unverified: those were verified by noise. Every simulated figure also shifts once, within its sampling error, because the random draws are new.

An exposed sensitive asset is compromised only when it is open

Section titled “An exposed sensitive asset is compromised only when it is open”

Affects you if a crown jewel of yours is itself internet-exposed - a public-subnet database, a VM with a public IP, a public bucket, a role anyone can assume.

Such an asset counted as compromised in every trial, which pinned the headline risk at 100% and zeroed every other fix’s risk reduction - while the path list showed nothing for it. Now:

  • Reachable (internet_exposed: a database behind a password, a VM): not compromised by exposure alone; it counts when an edge reaches it. It is reported by a new CRITICAL invariant, no-internet-exposed-sensitive-asset, on the Violations view.
  • Open to anyone (new property public_access: a bucket whose ACL lets anyone read, a role whose trust admits "*" without a condition): compromised as it stands, and now listed as a direct-access path - the asset alone, no steps, score 1, priority P1 - with a generated S3 public access block as its fix (a role gets a hint: its right trust policy names principals only you know). GraphQL: AttackPath.directAccess.

The Cloud Custodian collector now also treats an S3 grant to AuthenticatedUsers (any AWS account in the world) as public, and a write-only public grant as exposed but not open. An IAM trust admitting "*" under a Condition (e.g. aws:PrincipalOrgID) stays exposed but not open.

Action: none. Expect the headline risk to drop if one exposed asset was pinning it at 100%, a new CRITICAL violation for each merely reachable one, and a P1 path for each open one. Re-ingest Custodian and IAM output for the new property to appear.

Affects you if you run with SEED_IAM_USERS=true.

Only the path list started attacks from an identity whose credentials are assumed leaked; the risk simulation, the alternative routes (kShortestPaths) and the database path finder (ANALYZER_DB_PATHS=true) started from the internet alone, so that lens raised the path count and left the risk figure untouched. They now agree. Action: none; with the lens on, expect the risk figure to rise to include those routes.

The attacker-profile figures are anchored on each hop’s probability

Section titled “The attacker-profile figures are anchored on each hop’s probability”

Affects you if you read mixtureScore, posteriorMean, scoreCiLow/scoreCiHigh, profileScores, mixtureCompromiseProbability or profileCompromise.

The mixture treated a hop’s probability as the “criminal” profile’s, so with the default weak-attacker-heavy priors every figure was dragged below its inputs: a single hop at 0.9 read 0.78, and a two-hop CloudGoat route fell from 81% to 64% under a lens described as adding correlation. The profiles are now anchored so that, averaged, they give back each hop’s own probability: one hop reads exactly its probability, and a path reads between its independent score and its weakest hop. ATTACKER_PROFILE_PRIORS now changes the spread between profiles, not the level. Action: none; expect these figures to rise toward score. score, priority and the calibration grades are unaffected.

Narrower credible bands, and two properties that finally add up

Section titled “Narrower credible bands, and two properties that finally add up”

Affects you if you read sensitivityLow/sensitivityHigh, or your feeds set evidence_count or weight_cause on edges.

  • The risk figure’s credible band carried about ±4 points of sampling noise whatever the inputs; it now measures the inputs alone, so bands narrow where the evidence is strong.
  • evidence_count replaced the evidence a hop’s basis carries instead of adding to it: a KEV hop with one sighting came out less certain than a guess. It now adds.
  • Hops sharing a weight_cause count once, at the weakest, in the path score - as the risk simulation always sampled them - and set correlatedHops.
  • Alternative routes (kShortestPaths) now carry the interval, mixture, upper bound and join provenance a critical path does.

Action: none.

Ingest requests can be signed v2, and v1 can be turned off

Section titled “Ingest requests can be signed v2, and v1 can be turned off”

Affects you if anything signs ingest requests itself - a script built on the MANUAL recipe, a webhook, a CI step other than the gate.

The v1 signature (X-PerspectiveGraph-Signature: sha256=…) covers the body alone. The repository and commit a report counts against travel in the query, so a captured request could be replayed at any time, or its ?sha= changed to put its findings - or its lack of findings - on another commit. v2 (X-PerspectiveGraph-Signature-V2 with X-PerspectiveGraph-Timestamp) covers the time, method, path, parameters and body, and each signature is accepted once within five minutes. The gate subcommand, the GitHub Action and the Postman collection now send both.

Action: none to keep working - v1 stays accepted (INGEST_HMAC_ACCEPT_V1=true). To close the hole: move your senders to v2 (MANUAL, “Authentication”), watch perspectivegraph_ingest_signatures_total{version="v1"} stay at 0, then set INGEST_HMAC_ACCEPT_V1=false (Helm: ingest.hmacAcceptV1: false). A proxy in front of ingest must pass the path through unchanged: the path is signed.

Large events are split, and the gate waits for the whole report

Section titled “Large events are split, and the gate waits for the whole report”

Affects you if you ingest large cluster or account dumps, or run the merge gate.

An event over one bus message (1 MiB on a default NATS) used to be refused at ingest with a 502, so a large estate never reached the graph. It is now split into chunks. The ingest response carries a batch id, and GraphQL ingestBatch(id) says when every chunk has been applied; the gate waits for that and takes its verdict from a pass after it, because a pass between two chunks would have seen part of the report.

Action: none. A gate older than 1.20 still works against this engine, as before - without the wait.

An edge waiting for its endpoint is parked, not redelivered

Section titled “An edge waiting for its endpoint is parked, not redelivered”

Affects you if your feeds send edges to assets another feed describes.

Such an edge used to send its whole event back for redelivery, eight times in about four minutes, and then to the dead-letter stream. It is now parked and lands in the write that brings its endpoint, from any feed or replica, for up to seven days.

Action: none. Expect fewer dead-lettered events. perspectivegraph_graph_pending_edges counts what is parked, and the shipped PerspectiveGraphPendingEdgesGrowing alert fires when it keeps rising - a feed naming assets nothing else describes.

Reading a tenant no longer creates it; tenants share one connection pool

Section titled “Reading a tenant no longer creates it; tenants share one connection pool”

Affects you if you run several tenants, or sign users in with OIDC tenant claims.

A read under a tenant nobody had written used to create its graph, a pool of eight database connections and an analyzer loop - so the number of tenants, not the operator, set how many connections a replica needed. A never-written tenant now reads as an empty graph, and all tenants’ graphs share one pool of eight connections per replica.

Action: none. Size max_connections for eight graph connections per replica, whatever the tenant count (plus the governance pool and the leader’s connection).

The auth-denial alert watches a new counter

Section titled “The auth-denial alert watches a new counter”

Affects you if you loaded deploy/observability/prometheus-alerts.yaml.

PerspectiveGraphAuthDenialSpike matched code=~"401|403" on a counter that records only status classes (4xx), so it could never fire. It now reads perspectivegraph_auth_denied_total, which counts refused credentials on the API and on ingest by reason.

Action: reload the alert rules.

  • perspectivegraph healthz - the container healthcheck - speaks HTTPS when the API serves TLS itself; it used to mark such a backend unhealthy. It verifies the listener against the process’s own TLS_CERT_FILE and a name that certificate lists, so the certificate must carry a DNS name or IP address (any certificate a browser accepts does).
  • The file-backed governance stores force their writes to disk. On macOS, where a sync is slow, importing many verdicts into the file store takes a few seconds longer.
  • OIDC tokens signed ES256/ES384/ES512 are accepted, each only with a key of its curve.
  • A node listed twice in a graph - duplicate vertices written by concurrent replicas before 1.19 - counts once in the risk simulation. It used to count twice: a compromise probability of 2 and an interval of NaN.
  • A node written without a name no longer fails reading the whole graph on Apache AGE.

The event stream keeps what is still to be processed, and nothing else

Section titled “The event stream keeps what is still to be processed, and nothing else”

Affects you if you upgrade a deployment whose NATS already holds the PERSPECTIVE stream - which is every upgrade.

The stream used to keep every event ever ingested: limits retention with no limit set, on a volume nothing emptied. It now uses interest retention, so an event leaves as soon as the backend acknowledges it, and NATS_MAX_AGE (default 168h) drops one nobody drains. The dead-letter stream keeps its events for the same NATS_MAX_AGE. The backend updates both streams in place on its first start, and NATS then deletes every event already processed, so the disk that history held is freed at once; events still waiting are kept and handled as before, and nothing is replayed. Settings you made on the stream yourself - replicas, above all, on a clustered NATS - are now kept: the backend used to reset them at every start.

Action: none, with two exceptions.

  • A NATS older than 2.10 refuses to change a stream’s retention, and the backend stops at startup with stream configuration update can not change retention policy. Upgrade NATS (the chart and Compose ship 2.15), or delete the stream and let the backend recreate it.
  • A consumer of your own on the stream now holds events back: under interest retention an event stays until every consumer has acknowledged it. Remove consumers you no longer read from.

A listener that cannot start now stops the process

Section titled “A listener that cannot start now stops the process”

Affects you if your logs from the previous version contain http server failed.

That line meant the listener named in it - api, ingestion or metrics - never served, while the process stayed up without it. A taken port or an unreadable certificate now exits with status 1 and the reason on the last line, and so does a bus connection that closes for good or a consumer that stops.

Action: fix whatever the old log line named before upgrading, or the new version crash-loops on it - which is the point: the old one was silently missing a port.

Replicas no longer write the same asset twice

Section titled “Replicas no longer write the same asset twice”

Affects you if you ran more than one backend replica against Apache AGE - values-ha.yaml, or backend.replicas above 1.

AGE has no unique constraint, and replicas write concurrently, so two events naming the same asset at the same moment could each create a vertex for it. Writes to a graph now take a lock and converge on one vertex, but duplicates created before the upgrade stay: an update reaches all of them, so they never go stale and the TTL pruner never removes them.

Action: check each tenant’s graph (perspective is the default tenant’s; others are perspective_<tenant>):

LOAD 'age'; SET search_path = ag_catalog, "$user", public;
SELECT * FROM cypher('perspective', $$
MATCH (n) WITH n.id AS id, count(*) AS copies WHERE copies > 1 RETURN id, copies
$$) AS (id agtype, copies agtype);

No rows: nothing to do. Rows: the graph is derived from the feeds, so the clean fix is to rebuild it - take the backup (OPERATIONS §4), drop the graph (SELECT drop_graph('perspective', true);), restart the backend, and let the scanners and connectors re-ingest. Suppressions, tickets and the audit log are not in the graph and are not affected.

An edge waiting for its endpoint no longer holds back the rest of its event

Section titled “An edge waiting for its endpoint no longer holds back the rest of its event”

Affects you if your feeds send edges to assets another feed describes.

The graph refused such an edge until its endpoint arrived, and stopped the event there: the edges listed after it waited too, and after eight redeliveries (about four minutes) went to the dead-letter stream along with it. Now every node and every other edge is written, and only the waiting edges are retried. Action: none. Expect fewer dead-lettered events and routes that used to appear late, or not at all, to appear on the first pass.

A request may run at most twenty heavy analyses

Section titled “A request may run at most twenty heavy analyses”

Affects you if a script asks for many what-ifs in one request - typically verification across a whole remediationPlan, or attackPaths { remediations { verification } }.

Each fix’s verification, each whatIf, each riskSimulation with its own iterations or seed, and each kShortestPaths search is a full computation over the graph, and the query guard, which prices a document before it runs, cannot see how long a list will be. On a 4,344-node estate, verification across a 75-fix plan did not finish in five minutes. A request now runs the first twenty of them - identical ones count once - and the rest of its heavy fields answer with an error saying so. At most half the cores run them at once; a request that waits 20 s for one is told the server is busy.

Action: ask for one fix’s proof at a time with the new argument, remediationPlan(title: "…") { verification { … } }, or split the request.

AI answers need a signed-in caller, and have a rate limit of their own

Section titled “AI answers need a signed-in caller, and have a rate limit of their own”

Affects you if you publish a read-only instance (API_ANONYMOUS_ROLE=viewer) with an AI key configured, or several people share one client address.

/ai/* asked only for the viewer role, which a public instance gives every visitor - so the operator paid for anyone’s questions. Anonymous callers now get 403 whenever auth is on, and aiEnabled answers false to them, so the dashboard hides the AI features. Every caller is also limited by AI_RATE_PER_MIN (default 10 per client per minute).

Action: sign in to use the AI features on a published instance. If a team reaches the backend through one address, set TRUSTED_PROXY_CIDRS so each person is a client of their own, or raise AI_RATE_PER_MIN.


A Trivy scan of an image archive now reaches the merge gate

Section titled “A Trivy scan of an image archive now reaches the merge gate”

Affects you if you scan images from an archive - docker save, then trivy image --input image.tar - and run the merge gate.

Trivy reports an archive scan under its file path, and a path matches no workload. The image’s libraries and CVEs arrived joined to nothing, the commit still counted as analysed, and the gate answered clean on routes that ran through that image. The collector now names the image by the tag the archive was saved with, which Trivy records in the report, so those routes count.

Action: none - but expect pull requests that used to pass to fail. They were passing because the scan was never connected to the workload, not because the change was safe. An archive saved by image ID carries no tag and still cannot be joined; save it under the reference you deploy (docker save name:tag), or scan the image by reference.


The Kubernetes feed can now put a commit on the merge gate

Section titled “The Kubernetes feed can now put a commit on the merge gate”

Affects you if you post cluster dumps to /ingest/k8s with ?slug=&sha=, and you run the merge gate.

Those parameters used to be ignored by this collector: the gate blocks when a node on a path carries the commit, and only the scanner feeds stamped one - so a pull request that changed a manifest, which is how most routes open, could not turn the check red, while a dependency bump could. The dump now stamps the objects it contains, so those routes count.

Action: none, unless you were already sending those parameters. If you were, and the dump is a snapshot of the live cluster rather than what the pull request renders, the gate will start attributing the whole snapshot to that commit. Drop the parameters for snapshot feeds; keep them for helm template / kustomize build output of the commit under test. Objects the dump only references (cluster-admin, a ServiceAccount named by a binding) are never stamped.


The chart could not install at all, and now can

Section titled “The chart could not install at all, and now can”

Affects you if you ever tried helm install with the default values. It failed, and this release is the fix.

The bundled database and broker pods declared runAsNonRoot: true without a runAsUser, and every image the chart deployed leaves USER unset - which is root. The kubelet refuses that combination outright:

container has runAsNonRoot and image will run as root

Both pods sat in CreateContainerConfigError and the backend waited behind them in Init:0/2 forever. CI never saw it because it rendered the templates and checked them against the restricted Pod Security Standard - which they passed - and never installed them. make chart-install now stands up a kind cluster and installs the chart with default values on two Kubernetes versions, so this class of failure cannot return silently.

Action: none, if you were using the bundled database - it could not have been running, so there is nothing to migrate. An install pointed at your own PostgreSQL+AGE was never affected.

Affects you if you override the bundled database image, typically as --set postgres.image=... in a pipeline. It now follows the same shape as the backend and dashboard images:

postgres:
image:
repository: ghcr.io/luiacuaniello/perspectivegraph-postgres
tag: "" # empty = the chart's appVersion

A string value now fails to render rather than being ignored, which is the safe direction.

The bundled demo database is built here instead of pulled

Section titled “The bundled demo database is built here instead of pulled”

The image moves from apache/age:release_PG17_1.7.0 to ghcr.io/luiacuaniello/perspectivegraph-postgres, built from deploy/postgres/Dockerfile: the same PostgreSQL 17 and Apache AGE 1.7.0, on Alpine instead of Debian, signed with cosign and carrying an SBOM and provenance like the other two.

Action: none. The postgres uid is deliberately kept at 999, the Debian value, so an existing docker compose volume is read by the new image unchanged - this was tested by writing a graph with the old image and reading it back with the new one. On Kubernetes PGDATA moves to a subdirectory of the mount, which no running cluster can notice for the reason in the first note.

Why bother, for a demo: apache/age is not stale - it is the official postgres:17-trixie image plus the extension - but its Debian base carried fourteen criticals, thirteen of them perl and libxml2 with no fix published in any version, so no rebuild by anyone would have cleared them. Alpine ships no perl. The chart’s report goes from 430 findings to 4.

Same server, same version, no Linux userland around it - which was twenty of that image’s twenty-three findings. Action: none unless you run the compose stack with a custom health check for NATS: there is no shell in the image to run one, and docker-compose.yml now waits for the broker with a busybox container instead.


The chart refuses to publish an unauthenticated instance

Section titled “The chart refuses to publish an unauthenticated instance”

service.type is now a value, defaulting to ClusterIP, and LoadBalancer or NodePort is guarded exactly like the ingress: the chart refuses to render either without a credential. It is offered on purpose - without it, exposing the backend meant patching the Service by hand, which no guard in the chart could see. Nothing changes for an install that leaves it at ClusterIP.

The dashboard also carries a banner, not dismissible, whenever /auth/config reports that no credential is required. It is what covers the exposure a chart cannot see - a patched Service, a hand-written Ingress - and it appears in make demo too, which runs open by design.

Affects you if you install the Helm chart with ingress.enabled: true and have not configured a credential. helm upgrade will refuse to render rather than apply.

Enabling the ingress is the moment an install becomes reachable, and this chart’s ingress routes both /graphql — this environment’s map of how to breach it — and /ingest, the write side that decides what the engine reasons over. The backend has always refused to start unauthenticated under PG_ENV=production, but that gate only fires for an operator who declared production; an install left on the demo default was reachable and open, with nothing but a startup warning that scrolls past in a log.

The chart now fails to render in that combination. Set one of:

auth:
apiTokens: "s3cr3t:admin" # or oidc.jwksUrl with issuer and audience
ingest:
hmacSecret: "another-secret" # or hmacSecrets for per-tenant keys

Credentials supplied through secrets.existingSecret satisfy the guard: the chart cannot read a secret’s contents, so an operator using one is trusted rather than blocked.

If an open instance is the point — a public read-only demo — say so explicitly:

ingress:
allowUnauthenticated: true

Nothing changes for an install with ingress.enabled: false, which is the default, or for make demo and Docker Compose, which bind to 127.0.0.1 only.

The bundled demo database moves to PostgreSQL 17

Section titled “The bundled demo database moves to PostgreSQL 17”

Affects you if you run the bundled database - make demo, docker compose, or a Helm install left on postgres.enabled: true. An install pointed at your own PostgreSQL+AGE is unaffected, and that is what production should be doing anyway.

The image moves from apache/age:release_PG16_1.6.0 to release_PG17_1.7.0. A PostgreSQL major version cannot read the previous major’s data directory, so an existing demo volume will not start under it. The data is derived - the graph is rebuilt by re-ingesting - so the fix is to drop the volume:

Terminal window
make down # `docker compose down -v` removes the volumes
make demo

On Kubernetes, delete the PVC before upgrading if you were using the bundled database.

Why bother, for a demo: the older image carried 19 critical and 191 high advisories, and it is what Artifact Hub scans and reports on the chart’s page. The newer one is 14 and 97. Nothing there is in code this project ships - both first-party images scan clean - but a default install deploying it is a default install answering for it. Thirteen of the remaining critical findings have no fix available from Debian in any version.

The NATS image moves with it, from 2.12.11 to 2.14.6. That one was entirely ours to fix: all thirteen of its findings had upstream fixes, and the new image is clean of criticals. No action is needed - NATS reads no persistent state in this deployment.

The chart installs the app version it was built with, not latest

Section titled “The chart installs the app version it was built with, not latest”

Affects you if you install the Helm chart without setting backend.image.tag or frontend.image.tag - which is the default.

Both defaulted to latest. A chart is a versioned, signed artefact that was tested against one build of the application, and latest floats: chart 1.12.0 would deploy whatever image had been pushed most recently, which is not necessarily the one it declares. The default is now empty, and an empty tag resolves to the chart’s own appVersion.

Concretely, helm install with defaults moves from …/perspectivegraph:latest to …/perspectivegraph:v1.12.1. If you were relying on latest to pick up new images without touching your values, set it back explicitly:

backend:
image:
tag: latest

Better, pin a digest - tag: "@sha256:…" - which is what OPERATIONS asks for in production and what the release publishes.

There was a second, quieter consequence. Artifact Hub scans the images a chart deploys and publishes the report on the chart’s page; with a floating tag it was scanning something other than the release.

A security release. Three behaviours changed, each of them a control that now refuses something it used to allow. All three are silent in the sense that nothing crashes - so if one applies to you and you do not act, the effect is a feature quietly not working.

Affects you if GITHUB_TOKEN or GITLAB_TOKEN is set and you are not in dry-run.

PR comments, the merge-gate commit status and remediation PRs now write only to repositories you name. The destination was previously read from an ingested node property (repo_slug), which means it was chosen by whoever can post an event - and the ingest endpoint is reachable by every scanner holding the shared HMAC key. A success commit status in a repository where this check is required opens a merge gate, so this was worth closing at the cost of a required setting.

Set it to the repositories that are yours, as exact slugs or an owner wildcard:

Terminal window
REPO_ALLOWLIST=acme/payments-api,acme/*

With it empty, every real write is refused. You will see this once at start-up:

forge token set but REPO_ALLOWLIST is empty: every PR comment, commit status and
remediation PR will be refused

and one line per refusal, naming the repository it declined (pr comment refused: repository not allowed). POST /remediation/pr answers 422 naming the setting. Dry-run is exempt - it makes no outbound call - so a demo keeps printing what it would post.

POST /ingest/events rejects labels outside the ontology

Section titled “POST /ingest/events rejects labels outside the ontology”

Affects you if you hand-author events with a label or edge type that is not in the documented vocabulary, and you run the in-memory graph backend.

The vocabulary was already enforced by the Apache AGE store, so an AGE deployment saw no change; the in-memory backend accepted any string, and those values reached code that assumes a closed set - including the prompt the AI layer builds. The check moved to the ingest door and to the single writer into the graph.

A rejected request answers 400 and names the value:

outside the ontology: node "n": unknown label "MyCustomThing"

Map your events onto the labels and edge types listed in the manual. If you need a value that is not there, open an issue - adding one is a minor release, and a local string that only worked on one backend was never portable.

An apps-scoped principal can no longer act outside its applications

Section titled “An apps-scoped principal can no longer act outside its applications”

Affects you if any token or OIDC claim carries an apps allowlist. A principal without one is unaffected, and that is most deployments.

Suppressions, tickets and validations are keyed by attack-path id and were filtered by tenant alone, while the path reads behind them were already filtered by application. So a principal scoped to one application could suppress a path belonging to another - hiding a real finding from the team that owns it. Those boards are now filtered, and a write against a path outside the caller’s applications answers 404 attack path not found (or out of your scope).

If a scoped principal of yours legitimately needs a wider view, widen its apps claim; if it needs the whole tenant, drop the claim.

One thing deliberately did not change: the tenant-wide calibration and precision/recall aggregates are still tenant-wide. They measure the engine rather than any application, name no path or asset, and GraphQL serves the same numbers - so scoping only the REST board would have been a control in name only. It is written up in the threat model.