Upgrade notes
What changes for a running deployment when you move between versions - the settings you have to set, the requests that start failing, the behaviour that is no longer what it was.
This file exists because the CHANGELOG does not carry it. That file is generated from Conventional Commit subjects, so it names what changed in one line and drops the paragraph underneath explaining what an operator has to do about it. A one-line entry is fine for a bug fix and useless for a release that refuses writes until a new variable is set.
Only versions that need an action appear here. A version missing from this list is one you can upgrade into without touching your configuration - which is most of them. Follow the upgrade recipe in SUPPORT.md either way: pin by digest, take the backup, stage it.
1.25.0
Section titled “1.25.0”P1 means “act now”, and red-team verdicts move a route
Section titled “P1 means “act now”, and red-team verdicts move a route”Affects you if you triage by band (P1/P2/P3), read priority or priorityLabel from
the API or the MCP tools, or alert on P1 counts.
A route with a runtime alert into AdministratorAccess used to read P2: the blended priority weighs exploitability at a third and could not reach P1 without a KEV entry. Now:
- A fact about the route makes it P1: a runtime alert on it, an asset open to anyone, a
KEV weakness on a route an attacker is likely to complete, or a likely route (≥ 80%, on
evidence rather than default weights) into a high-value asset.
priorityReason(new) says which. Every P1 path sits at 70 plus 30% of its blended priority, so the band keeps an order and a natural P1 of 76 now reads 92.8. - Red-team and BAS verdicts re-band a route where paths are listed
(
attackPaths, the dashboard, the MCP tools): confirmed → P1; refuted → P3, unless a runtime alert or open access contradicts the test - then it stays and says the evidence conflicts. The Priority recorded with a verdict, which grades the triage order, still contains none. - Expect more P1s where routes are runtime-confirmed or tested; expect refuted routes at the bottom. The dashboard no longer re-sorts refuted routes itself.
Corrected ATT&CK mapping
Section titled “Corrected ATT&CK mapping”AFFECTS (a library has a CVE) and DEPENDS_ON (an image ships a library) no longer carry a
technique: they are facts about software, not attacker actions. An exploit is T1190
(initial access) until the attacker is inside, then T1210 (lateral movement), and the
exposure before an exploit carries no technique of its own. CAN_ESCALATE_TO is
T1098.003 (Account Manipulation: Additional Cloud Roles), no longer T1078.004. Anything
that stored or keyed on the old technique ids per hop will see different ones.
Calibration reads that could not be trusted
Section titled “Calibration reads that could not be trusted”brier_recalibrated is measured out of sample at every size - leave-one-out below 20
outcomes, where it was the in-sample fit and read a perfect 0.000 on a dozen points. Below 30
outcomes the Trust page headlines “Not enough outcomes yet” with the direction beside it, and
an AUC below its floor shows the counts it rests on instead of an interval.
Action: none; nothing to configure. Check anything that alerts on P1 counts or keys on per-hop ATT&CK ids.
1.24.0
Section titled “1.24.0”The chart’s bus authenticates and keeps its stream, and its database password is no longer shared
Section titled “The chart’s bus authenticates and keeps its stream, and its database password is no longer shared”Affects you if you install with the Helm chart.
The chart’s NATS accepted any client, so any pod that could reach it published events straight
into the graph - past the ingest webhook’s signature check - or deleted the stream. It kept that
stream on an emptyDir, so a restart lost every queued event. And the bundled Postgres used the
password perspective on every install.
What changes on helm upgrade:
- NATS requires a user and password. The chart generates the password into a Secret of its
own (
<release>-perspectivegraph-nats-auth), gives it to the backend, and keeps it across upgrades. NATS and the backend both roll during the upgrade, and ingest pauses until both are up. An external NATS is untouched unless you setnats.auth.userandnats.auth.password(ornats.auth.existingSecret); the backend readsNATS_USER/NATS_PASSWORDin any deployment. - NATS is a StatefulSet with a 2 GiB volume (
nats.persistence). The upgrade replaces the Deployment, so events still queued in the old pod are lost this once - as every NATS restart lost them before. The volume needs a default StorageClass, as the bundled Postgres’s already does;nats.persistence.enabled: falsekeeps the emptyDir. postgres.auth.passworddefaults to empty. A new install gets a random password. An existing one keeps the password in its Secret -perspectiveunless you set one - because Postgres reads it only when it first creates its data directory. To rotate it,ALTER ROLEin the database, then setpostgres.auth.password.secrets.existingSecretworks for the bundled Postgres. Its pod named the chart’s own Secret, which does not exist when you bring yours, so it could not start. It now reads yours, as the backend always did.networkPolicy: only the backend may reach the bundled NATS and Postgres. Off by default, on invalues-production.yaml;backendFromalso closes the backend to all but the dashboard and the peers you name.
Action if you render with helm template (Argo CD, Flux): a render without the cluster
cannot read the existing Secrets, so it would draw new passwords on every sync - and a new
Postgres password locks the backend out of the database it already has. Before upgrading, set
postgres.auth.password to the current one (perspective if you never set it) and
nats.auth.password, or bring both through secrets.existingSecret and
nats.auth.existingSecret. An external Postgres now needs its password set explicitly there
too.
Action to keep the old behaviour (not recommended): nats.auth.enabled: false,
nats.persistence.enabled: false.
1.23.0
Section titled “1.23.0”The graph forgets what its sources stop reporting
Section titled “The graph forgets what its sources stop reporting”Affects you if you ingest Trivy or Semgrep reports, run the AWS connector, or rely on something staying in the graph after its source stopped listing it.
Until now the only way out of the graph was GRAPH_TTL, off by default, so a CVE fixed in an
image, an instance terminated or a role deleted stayed - with the attack paths through it -
until someone pruned. The engine now records who asserted each node and edge. When a source
that describes a scope in full sends that scope again, whatever it said before and no longer
says is withdrawn, once the whole ingest has landed. An element leaves the graph when no source
asserts it any more.
What changes:
- Trivy scans of an image by reference, Semgrep reports with
?repo=, and the AWS connector (per account; the network per account and region) are complete snapshots without any change on your side. A fixed CVE disappears at the next scan. - Everything else is unchanged unless you say so with
?snapshot=<scope>on the ingest webhook -?snapshot=cluster:prodon a dump of the whole cluster. Name only what you really send in full: what the scope omits is removed. - A Trivy scan filtered by severity or fixability is complete for its filter only. If
one pipeline sends the same image filtered and another unfiltered, they now undo each
other’s findings: send the filtered one with
?snapshot=none. - A pull request’s scan never removes anything, and
?snapshot=together withslug/sha/pris refused with400. - Elements already in the graph when you upgrade are never removed this way - nothing
recorded who sent them - only by
GRAPH_TTL. To clean up leftovers from before, enableGRAPH_TTLfor a cycle, or re-create the graph and let the feeds rebuild it. - The Postgres graph gains a provenance table beside the parked edges
(
<graph>._pg_provenance, plus_pg_removals). It is created and filled on first write, under the graph’s write lock, and needs no migration step. Every write now also records its origin, one row per element per source. - With
ANALYZER_INCREMENTAL=true, the analyzer re-reads the whole graph after any removal, on every replica. - The Helm ingress now raises the nginx-based controllers’ 1 MiB request-body limit to the
32 MiB the backend accepts, for
/ingestand/gate(ingress.maxBodySize). A body-size annotation you set yourself wins.
Action to keep the old behaviour: GRAPH_SWEEP=false (Helm graph.sweep: false): nothing
is removed except by GRAPH_TTL.
1.22.0
Section titled “1.22.0”The merge gate counts what the change adds, and writes nothing
Section titled “The merge gate counts what the change adds, and writes nothing”Affects you if you run the merge gate - the GitHub Action, perspectivegraph gate, or
the Trivy plugin.
The gate used to count every critical path through an asset stamped with the commit. A pull
request that rescanned an image already in production, or re-rendered a deployment’s
manifests, carried the commit onto assets that were on routes long before it - and every one
of those routes blocked it: in the demo lab, nine paths where the change itself opened two.
It now applies the report to a copy of the estate and counts only the routes the change
opens or makes likelier. The routes it merely touches are reported (preexisting) and
do not count. Nothing is written: the comparison runs on POST /gate/impact, on the API
port, with the bearer token.
What changes for a running pipeline:
- Fewer red checks, by design: the ones that remain are routes the change caused.
max-criticalnow counts those. - Server mode no longer ingests the report. The live graph does not record a pull
request’s scan unless you ask:
persist: true(CLI-persist) sends it to the webhook after the verdict. The engine’s own pull-request comments and commit status are driven by what is ingested, so they needpersist- or the webhook step you already had. - Server mode needs
tokenfor the comparison if API auth is on (the comparison is an authenticated API call, refused to anonymous callers on a public instance).ingestandhmac-secretare needed only withpersistorattribution: commit. - Local mode compares with the estate you give it. Add
base-reports- the scan of what runs now - or the routes through the scanned image count as the change’s, since the estate knows none of its findings. - A proxy in front of the API must pass
/gate/- the bundled dashboard nginx and the Helm ingress do - and let a report through: raise an ingress controller’s 1 MiB body limit (ingress-nginx:nginx.ingress.kubernetes.io/proxy-body-size: 32m), as/ingestalready needed.
Action to keep the old behaviour: attribution: commit (CLI -attribution commit). The
gate also falls back to it by itself - saying so in the log and in the attribution output -
when there is no report to compare, or the engine predates 1.22. A gate binary older than
1.22 under the new action runs per commit with a warning. prVerdict is unchanged.
The engine’s commit status and PR comments count what the gate counts
Section titled “The engine’s commit status and PR comments count what the gate counts”Affects you if the engine itself posts to your pull requests - GITHUB_TOKEN or
GITLAB_TOKEN set, with REPO_ALLOWLIST.
The commit status perspectivegraph/attack-paths and the PR/MR comments counted every critical
path through an asset stamped with the commit, like the gate before this release. They now
apply the gate’s rule, comparing each analysis pass with the one before it. A route that is new,
or likelier, belongs to the pull-request commits on it that arrived in between. A route that was
already there when a commit arrived gets no comment and does not turn its status red, and the
status description says how many there were. A route that appears later through a commit’s
assets belongs to the pull request that arrived with it. If none did, it counts against the
commits already on it, as “since this change arrived”.
What changes:
- Fewer red statuses and fewer comments, by design, and the words change: “N critical attack path(s) opened or made likelier by this change”. A commit already in the graph when the engine starts - after a restart or an upgrade - has no “before”: every route through it counts, and the status says “in the graph before the engine was watching”. That is the gate’s rule for a commit the engine already holds.
- Every commit on a route is judged, not only the first one found on it. Under the new rule the first could be innocent and the route belong to the second.
- A status is posted when it changes, not re-posted on every analysis pass. GitHub keeps at most 1,000 statuses per commit and context.
- Fixed: with more than one tenant, one tenant’s analysis pass posted
successon another tenant’s red commits, and the last route of a tenant closing never cleared its red status.
Action to keep the old behaviour: PR_ATTRIBUTION=commit (Helm prAttribution: commit),
which restores the old rule and wording. The three changes above stay.
1.21.0
Section titled “1.21.0”Every probability the engine prints is recomputed under corrected rules, so the dashboard’s numbers move once on upgrade. Nothing needs configuring; what follows is what moves and why, so a step in a trend line is not mistaken for a change in the estate.
A fix is “verified” only when it protects something
Section titled “A fix is “verified” only when it protects something”Affects you if you read the verified badge on remediations, verification over
GraphQL, or whatIf results.
A fix’s verification compared two simulations of 800 trials that drew their random numbers independently, and called the fix verified when they differed by 0.05 points - against a noise of about ±2.5. A fix cutting an edge that leads nowhere was verified about half the time, and a what-if could show the risk rising after a cut. Both simulations now run the same trials, so the difference is the cut’s own effect: never negative, and exactly zero when the cut protects nothing.
Verification also reads a new measure, expectedReduction - the drop in the expected
number of sensitive assets compromised - on verification and on whatIf. The old
riskReductionPct / riskReduction (P(any asset compromised)) stays, but it saturates:
while one asset is compromised in every trial it reads 0 for every other fix.
Action: none. Expect some fixes to turn from verified to unverified: those were
verified by noise. Every simulated figure also shifts once, within its sampling error,
because the random draws are new.
An exposed sensitive asset is compromised only when it is open
Section titled “An exposed sensitive asset is compromised only when it is open”Affects you if a crown jewel of yours is itself internet-exposed - a public-subnet database, a VM with a public IP, a public bucket, a role anyone can assume.
Such an asset counted as compromised in every trial, which pinned the headline risk at 100% and zeroed every other fix’s risk reduction - while the path list showed nothing for it. Now:
- Reachable (
internet_exposed: a database behind a password, a VM): not compromised by exposure alone; it counts when an edge reaches it. It is reported by a new CRITICAL invariant,no-internet-exposed-sensitive-asset, on the Violations view. - Open to anyone (new property
public_access: a bucket whose ACL lets anyone read, a role whose trust admits"*"without a condition): compromised as it stands, and now listed as a direct-access path - the asset alone, no steps, score 1, priority P1 - with a generated S3 public access block as its fix (a role gets a hint: its right trust policy names principals only you know). GraphQL:AttackPath.directAccess.
The Cloud Custodian collector now also treats an S3 grant to AuthenticatedUsers (any AWS
account in the world) as public, and a write-only public grant as exposed but not open. An
IAM trust admitting "*" under a Condition (e.g. aws:PrincipalOrgID) stays exposed but
not open.
Action: none. Expect the headline risk to drop if one exposed asset was pinning it at 100%, a new CRITICAL violation for each merely reachable one, and a P1 path for each open one. Re-ingest Custodian and IAM output for the new property to appear.
Credential-origin seeds count everywhere
Section titled “Credential-origin seeds count everywhere”Affects you if you run with SEED_IAM_USERS=true.
Only the path list started attacks from an identity whose credentials are assumed leaked;
the risk simulation, the alternative routes (kShortestPaths) and the database path finder
(ANALYZER_DB_PATHS=true) started from the internet alone, so that lens raised the path
count and left the risk figure untouched. They now agree. Action: none; with the lens
on, expect the risk figure to rise to include those routes.
The attacker-profile figures are anchored on each hop’s probability
Section titled “The attacker-profile figures are anchored on each hop’s probability”Affects you if you read mixtureScore, posteriorMean, scoreCiLow/scoreCiHigh,
profileScores, mixtureCompromiseProbability or profileCompromise.
The mixture treated a hop’s probability as the “criminal” profile’s, so with the default
weak-attacker-heavy priors every figure was dragged below its inputs: a single hop at 0.9
read 0.78, and a two-hop CloudGoat route fell from 81% to 64% under a lens described as
adding correlation. The profiles are now anchored so that, averaged, they give back each
hop’s own probability: one hop reads exactly its probability, and a path reads between its
independent score and its weakest hop. ATTACKER_PROFILE_PRIORS now changes the spread
between profiles, not the level. Action: none; expect these figures to rise toward
score. score, priority and the calibration grades are unaffected.
Narrower credible bands, and two properties that finally add up
Section titled “Narrower credible bands, and two properties that finally add up”Affects you if you read sensitivityLow/sensitivityHigh, or your feeds set
evidence_count or weight_cause on edges.
- The risk figure’s credible band carried about ±4 points of sampling noise whatever the inputs; it now measures the inputs alone, so bands narrow where the evidence is strong.
evidence_countreplaced the evidence a hop’s basis carries instead of adding to it: a KEV hop with one sighting came out less certain than a guess. It now adds.- Hops sharing a
weight_causecount once, at the weakest, in the path score - as the risk simulation always sampled them - and setcorrelatedHops. - Alternative routes (
kShortestPaths) now carry the interval, mixture, upper bound and join provenance a critical path does.
Action: none.
1.20.0
Section titled “1.20.0”Ingest requests can be signed v2, and v1 can be turned off
Section titled “Ingest requests can be signed v2, and v1 can be turned off”Affects you if anything signs ingest requests itself - a script built on the MANUAL recipe, a webhook, a CI step other than the gate.
The v1 signature (X-PerspectiveGraph-Signature: sha256=…) covers the body alone. The
repository and commit a report counts against travel in the query, so a captured request
could be replayed at any time, or its ?sha= changed to put its findings - or its lack of
findings - on another commit. v2 (X-PerspectiveGraph-Signature-V2 with
X-PerspectiveGraph-Timestamp) covers the time, method, path, parameters and body, and
each signature is accepted once within five minutes. The gate subcommand, the GitHub
Action and the Postman collection now send both.
Action: none to keep working - v1 stays accepted (INGEST_HMAC_ACCEPT_V1=true). To
close the hole: move your senders to v2 (MANUAL, “Authentication”), watch
perspectivegraph_ingest_signatures_total{version="v1"} stay at 0, then set
INGEST_HMAC_ACCEPT_V1=false (Helm: ingest.hmacAcceptV1: false). A proxy in front of
ingest must pass the path through unchanged: the path is signed.
Large events are split, and the gate waits for the whole report
Section titled “Large events are split, and the gate waits for the whole report”Affects you if you ingest large cluster or account dumps, or run the merge gate.
An event over one bus message (1 MiB on a default NATS) used to be refused at ingest with
a 502, so a large estate never reached the graph. It is now split into chunks. The ingest
response carries a batch id, and GraphQL ingestBatch(id) says when every chunk has
been applied; the gate waits for that and takes its verdict from a pass after it, because
a pass between two chunks would have seen part of the report.
Action: none. A gate older than 1.20 still works against this engine, as before - without the wait.
An edge waiting for its endpoint is parked, not redelivered
Section titled “An edge waiting for its endpoint is parked, not redelivered”Affects you if your feeds send edges to assets another feed describes.
Such an edge used to send its whole event back for redelivery, eight times in about four minutes, and then to the dead-letter stream. It is now parked and lands in the write that brings its endpoint, from any feed or replica, for up to seven days.
Action: none. Expect fewer dead-lettered events. perspectivegraph_graph_pending_edges
counts what is parked, and the shipped PerspectiveGraphPendingEdgesGrowing alert fires
when it keeps rising - a feed naming assets nothing else describes.
Reading a tenant no longer creates it; tenants share one connection pool
Section titled “Reading a tenant no longer creates it; tenants share one connection pool”Affects you if you run several tenants, or sign users in with OIDC tenant claims.
A read under a tenant nobody had written used to create its graph, a pool of eight database connections and an analyzer loop - so the number of tenants, not the operator, set how many connections a replica needed. A never-written tenant now reads as an empty graph, and all tenants’ graphs share one pool of eight connections per replica.
Action: none. Size max_connections for eight graph connections per replica,
whatever the tenant count (plus the governance pool and the leader’s connection).
The auth-denial alert watches a new counter
Section titled “The auth-denial alert watches a new counter”Affects you if you loaded deploy/observability/prometheus-alerts.yaml.
PerspectiveGraphAuthDenialSpike matched code=~"401|403" on a counter that records only
status classes (4xx), so it could never fire. It now reads
perspectivegraph_auth_denied_total, which counts refused credentials on the API and on
ingest by reason.
Action: reload the alert rules.
Smaller changes, no action
Section titled “Smaller changes, no action”perspectivegraph healthz- the container healthcheck - speaks HTTPS when the API serves TLS itself; it used to mark such a backend unhealthy. It verifies the listener against the process’s ownTLS_CERT_FILEand a name that certificate lists, so the certificate must carry a DNS name or IP address (any certificate a browser accepts does).- The file-backed governance stores force their writes to disk. On macOS, where a sync is slow, importing many verdicts into the file store takes a few seconds longer.
- OIDC tokens signed ES256/ES384/ES512 are accepted, each only with a key of its curve.
- A node listed twice in a graph - duplicate vertices written by concurrent replicas before 1.19 - counts once in the risk simulation. It used to count twice: a compromise probability of 2 and an interval of NaN.
- A node written without a name no longer fails reading the whole graph on Apache AGE.
1.19.0
Section titled “1.19.0”The event stream keeps what is still to be processed, and nothing else
Section titled “The event stream keeps what is still to be processed, and nothing else”Affects you if you upgrade a deployment whose NATS already holds the PERSPECTIVE
stream - which is every upgrade.
The stream used to keep every event ever ingested: limits retention with no limit set, on
a volume nothing emptied. It now uses interest retention, so an event leaves as soon as the
backend acknowledges it, and NATS_MAX_AGE (default 168h) drops one nobody drains. The
dead-letter stream keeps its events for the same NATS_MAX_AGE. The backend updates both
streams in place on its first start, and NATS then deletes every event already processed,
so the disk that history held is freed at once; events still waiting are kept and handled
as before, and nothing is replayed. Settings you made on the stream yourself - replicas,
above all, on a clustered NATS - are now kept: the backend used to reset them at every
start.
Action: none, with two exceptions.
- A NATS older than 2.10 refuses to change a stream’s retention, and the backend stops
at startup with
stream configuration update can not change retention policy. Upgrade NATS (the chart and Compose ship 2.15), or delete the stream and let the backend recreate it. - A consumer of your own on the stream now holds events back: under interest retention an event stays until every consumer has acknowledged it. Remove consumers you no longer read from.
A listener that cannot start now stops the process
Section titled “A listener that cannot start now stops the process”Affects you if your logs from the previous version contain http server failed.
That line meant the listener named in it - api, ingestion or metrics - never served,
while the process stayed up without it. A taken port or an unreadable certificate now exits
with status 1 and the reason on the last line, and so does a bus connection that closes
for good or a consumer that stops.
Action: fix whatever the old log line named before upgrading, or the new version crash-loops on it - which is the point: the old one was silently missing a port.
Replicas no longer write the same asset twice
Section titled “Replicas no longer write the same asset twice”Affects you if you ran more than one backend replica against Apache AGE -
values-ha.yaml, or backend.replicas above 1.
AGE has no unique constraint, and replicas write concurrently, so two events naming the same asset at the same moment could each create a vertex for it. Writes to a graph now take a lock and converge on one vertex, but duplicates created before the upgrade stay: an update reaches all of them, so they never go stale and the TTL pruner never removes them.
Action: check each tenant’s graph (perspective is the default tenant’s; others are
perspective_<tenant>):
LOAD 'age'; SET search_path = ag_catalog, "$user", public;SELECT * FROM cypher('perspective', $$ MATCH (n) WITH n.id AS id, count(*) AS copies WHERE copies > 1 RETURN id, copies$$) AS (id agtype, copies agtype);No rows: nothing to do. Rows: the graph is derived from the feeds, so the clean fix is to
rebuild it - take the backup (OPERATIONS §4), drop the graph
(SELECT drop_graph('perspective', true);), restart the backend, and let the scanners and
connectors re-ingest. Suppressions, tickets and the audit log are not in the graph and are
not affected.
An edge waiting for its endpoint no longer holds back the rest of its event
Section titled “An edge waiting for its endpoint no longer holds back the rest of its event”Affects you if your feeds send edges to assets another feed describes.
The graph refused such an edge until its endpoint arrived, and stopped the event there: the edges listed after it waited too, and after eight redeliveries (about four minutes) went to the dead-letter stream along with it. Now every node and every other edge is written, and only the waiting edges are retried. Action: none. Expect fewer dead-lettered events and routes that used to appear late, or not at all, to appear on the first pass.
A request may run at most twenty heavy analyses
Section titled “A request may run at most twenty heavy analyses”Affects you if a script asks for many what-ifs in one request - typically
verification across a whole remediationPlan, or
attackPaths { remediations { verification } }.
Each fix’s verification, each whatIf, each riskSimulation with its own iterations or
seed, and each kShortestPaths search is a full computation over the graph, and the
query guard, which prices a document before it runs, cannot see how long a list will be. On
a 4,344-node estate, verification across a 75-fix plan did not finish in five minutes. A
request now runs the first twenty of them - identical ones count once - and the rest of
its heavy fields answer with an error saying so. At most half the cores run them at once;
a request that waits 20 s for one is told the server is busy.
Action: ask for one fix’s proof at a time with the new argument,
remediationPlan(title: "…") { verification { … } }, or split the request.
AI answers need a signed-in caller, and have a rate limit of their own
Section titled “AI answers need a signed-in caller, and have a rate limit of their own”Affects you if you publish a read-only instance (API_ANONYMOUS_ROLE=viewer) with an
AI key configured, or several people share one client address.
/ai/* asked only for the viewer role, which a public instance gives every visitor - so
the operator paid for anyone’s questions. Anonymous callers now get 403 whenever auth is on,
and aiEnabled answers false to them, so the dashboard hides the AI features. Every
caller is also limited by AI_RATE_PER_MIN (default 10 per client per minute).
Action: sign in to use the AI features on a published instance. If a team reaches the
backend through one address, set TRUSTED_PROXY_CIDRS so each person is a client of their
own, or raise AI_RATE_PER_MIN.
1.18.0
Section titled “1.18.0”A Trivy scan of an image archive now reaches the merge gate
Section titled “A Trivy scan of an image archive now reaches the merge gate”Affects you if you scan images from an archive - docker save, then
trivy image --input image.tar - and run the merge gate.
Trivy reports an archive scan under its file path, and a path matches no workload. The image’s libraries and CVEs arrived joined to nothing, the commit still counted as analysed, and the gate answered clean on routes that ran through that image. The collector now names the image by the tag the archive was saved with, which Trivy records in the report, so those routes count.
Action: none - but expect pull requests that used to pass to fail. They were passing
because the scan was never connected to the workload, not because the change was safe.
An archive saved by image ID carries no tag and still cannot be joined; save it under the
reference you deploy (docker save name:tag), or scan the image by reference.
1.17.0
Section titled “1.17.0”The Kubernetes feed can now put a commit on the merge gate
Section titled “The Kubernetes feed can now put a commit on the merge gate”Affects you if you post cluster dumps to /ingest/k8s with ?slug=&sha=, and you
run the merge gate.
Those parameters used to be ignored by this collector: the gate blocks when a node on a path carries the commit, and only the scanner feeds stamped one - so a pull request that changed a manifest, which is how most routes open, could not turn the check red, while a dependency bump could. The dump now stamps the objects it contains, so those routes count.
Action: none, unless you were already sending those parameters. If you were, and the
dump is a snapshot of the live cluster rather than what the pull request renders, the gate
will start attributing the whole snapshot to that commit. Drop the parameters for snapshot
feeds; keep them for helm template / kustomize build output of the commit under test.
Objects the dump only references (cluster-admin, a ServiceAccount named by a binding)
are never stamped.
1.12.7
Section titled “1.12.7”The chart could not install at all, and now can
Section titled “The chart could not install at all, and now can”Affects you if you ever tried helm install with the default values. It failed, and
this release is the fix.
The bundled database and broker pods declared runAsNonRoot: true without a runAsUser,
and every image the chart deployed leaves USER unset - which is root. The kubelet refuses
that combination outright:
container has runAsNonRoot and image will run as rootBoth pods sat in CreateContainerConfigError and the backend waited behind them in
Init:0/2 forever. CI never saw it because it rendered the templates and checked them
against the restricted Pod Security Standard - which they passed - and never installed
them. make chart-install now stands up a kind cluster and installs the chart with default
values on two Kubernetes versions, so this class of failure cannot return silently.
Action: none, if you were using the bundled database - it could not have been running, so there is nothing to migrate. An install pointed at your own PostgreSQL+AGE was never affected.
postgres.image is now a map, not a string
Section titled “postgres.image is now a map, not a string”Affects you if you override the bundled database image, typically as
--set postgres.image=... in a pipeline. It now follows the same shape as the backend and
dashboard images:
postgres: image: repository: ghcr.io/luiacuaniello/perspectivegraph-postgres tag: "" # empty = the chart's appVersionA string value now fails to render rather than being ignored, which is the safe direction.
The bundled demo database is built here instead of pulled
Section titled “The bundled demo database is built here instead of pulled”The image moves from apache/age:release_PG17_1.7.0 to
ghcr.io/luiacuaniello/perspectivegraph-postgres, built from deploy/postgres/Dockerfile:
the same PostgreSQL 17 and Apache AGE 1.7.0, on Alpine instead of Debian, signed with
cosign and carrying an SBOM and provenance like the other two.
Action: none. The postgres uid is deliberately kept at 999, the Debian value, so an
existing docker compose volume is read by the new image unchanged - this was tested by
writing a graph with the old image and reading it back with the new one. On Kubernetes
PGDATA moves to a subdirectory of the mount, which no running cluster can notice for the
reason in the first note.
Why bother, for a demo: apache/age is not stale - it is the official postgres:17-trixie
image plus the extension - but its Debian base carried fourteen criticals, thirteen of
them perl and libxml2 with no fix published in any version, so no rebuild by anyone would
have cleared them. Alpine ships no perl. The chart’s report goes from 430 findings to 4.
NATS moves to the scratch image
Section titled “NATS moves to the scratch image”Same server, same version, no Linux userland around it - which was twenty of that image’s
twenty-three findings. Action: none unless you run the compose stack with a custom
health check for NATS: there is no shell in the image to run one, and docker-compose.yml
now waits for the broker with a busybox container instead.
1.12.5
Section titled “1.12.5”The chart refuses to publish an unauthenticated instance
Section titled “The chart refuses to publish an unauthenticated instance”service.type is now a value, defaulting to ClusterIP, and LoadBalancer or NodePort
is guarded exactly like the ingress: the chart refuses to render either without a
credential. It is offered on purpose - without it, exposing the backend meant patching the
Service by hand, which no guard in the chart could see. Nothing changes for an install
that leaves it at ClusterIP.
The dashboard also carries a banner, not dismissible, whenever /auth/config reports that
no credential is required. It is what covers the exposure a chart cannot see - a patched
Service, a hand-written Ingress - and it appears in make demo too, which runs open by
design.
Affects you if you install the Helm chart with ingress.enabled: true and have not
configured a credential. helm upgrade will refuse to render rather than apply.
Enabling the ingress is the moment an install becomes reachable, and this chart’s ingress
routes both /graphql — this environment’s map of how to breach it — and /ingest, the
write side that decides what the engine reasons over. The backend has always refused to
start unauthenticated under PG_ENV=production, but that gate only fires for an operator
who declared production; an install left on the demo default was reachable and open, with
nothing but a startup warning that scrolls past in a log.
The chart now fails to render in that combination. Set one of:
auth: apiTokens: "s3cr3t:admin" # or oidc.jwksUrl with issuer and audienceingest: hmacSecret: "another-secret" # or hmacSecrets for per-tenant keysCredentials supplied through secrets.existingSecret satisfy the guard: the chart cannot
read a secret’s contents, so an operator using one is trusted rather than blocked.
If an open instance is the point — a public read-only demo — say so explicitly:
ingress: allowUnauthenticated: trueNothing changes for an install with ingress.enabled: false, which is the default, or for
make demo and Docker Compose, which bind to 127.0.0.1 only.
1.12.4
Section titled “1.12.4”The bundled demo database moves to PostgreSQL 17
Section titled “The bundled demo database moves to PostgreSQL 17”Affects you if you run the bundled database - make demo, docker compose, or a Helm
install left on postgres.enabled: true. An install pointed at your own PostgreSQL+AGE is
unaffected, and that is what production should be doing anyway.
The image moves from apache/age:release_PG16_1.6.0 to release_PG17_1.7.0. A
PostgreSQL major version cannot read the previous major’s data directory, so an existing
demo volume will not start under it. The data is derived - the graph is rebuilt by
re-ingesting - so the fix is to drop the volume:
make down # `docker compose down -v` removes the volumesmake demoOn Kubernetes, delete the PVC before upgrading if you were using the bundled database.
Why bother, for a demo: the older image carried 19 critical and 191 high advisories, and it is what Artifact Hub scans and reports on the chart’s page. The newer one is 14 and 97. Nothing there is in code this project ships - both first-party images scan clean - but a default install deploying it is a default install answering for it. Thirteen of the remaining critical findings have no fix available from Debian in any version.
The NATS image moves with it, from 2.12.11 to 2.14.6. That one was entirely ours to fix: all thirteen of its findings had upstream fixes, and the new image is clean of criticals. No action is needed - NATS reads no persistent state in this deployment.
1.12.1
Section titled “1.12.1”The chart installs the app version it was built with, not latest
Section titled “The chart installs the app version it was built with, not latest”Affects you if you install the Helm chart without setting backend.image.tag or
frontend.image.tag - which is the default.
Both defaulted to latest. A chart is a versioned, signed artefact that was tested
against one build of the application, and latest floats: chart 1.12.0 would deploy
whatever image had been pushed most recently, which is not necessarily the one it
declares. The default is now empty, and an empty tag resolves to the chart’s own
appVersion.
Concretely, helm install with defaults moves from …/perspectivegraph:latest to
…/perspectivegraph:v1.12.1. If you were relying on latest to pick up new images
without touching your values, set it back explicitly:
backend: image: tag: latestBetter, pin a digest - tag: "@sha256:…" - which is what
OPERATIONS asks for in production and what the release publishes.
There was a second, quieter consequence. Artifact Hub scans the images a chart deploys and publishes the report on the chart’s page; with a floating tag it was scanning something other than the release.
1.11.2
Section titled “1.11.2”A security release. Three behaviours changed, each of them a control that now refuses something it used to allow. All three are silent in the sense that nothing crashes - so if one applies to you and you do not act, the effect is a feature quietly not working.
Forge writes need REPO_ALLOWLIST
Section titled “Forge writes need REPO_ALLOWLIST”Affects you if GITHUB_TOKEN or GITLAB_TOKEN is set and you are not in dry-run.
PR comments, the merge-gate commit status and remediation PRs now write only to
repositories you name. The destination was previously read from an ingested node property
(repo_slug), which means it was chosen by whoever can post an event - and the ingest
endpoint is reachable by every scanner holding the shared HMAC key. A success commit
status in a repository where this check is required opens a merge gate, so this was worth
closing at the cost of a required setting.
Set it to the repositories that are yours, as exact slugs or an owner wildcard:
REPO_ALLOWLIST=acme/payments-api,acme/*With it empty, every real write is refused. You will see this once at start-up:
forge token set but REPO_ALLOWLIST is empty: every PR comment, commit status andremediation PR will be refusedand one line per refusal, naming the repository it declined (pr comment refused: repository not allowed). POST /remediation/pr answers 422 naming the setting. Dry-run
is exempt - it makes no outbound call - so a demo keeps printing what it would post.
POST /ingest/events rejects labels outside the ontology
Section titled “POST /ingest/events rejects labels outside the ontology”Affects you if you hand-author events with a label or edge type that is not in the documented vocabulary, and you run the in-memory graph backend.
The vocabulary was already enforced by the Apache AGE store, so an AGE deployment saw no change; the in-memory backend accepted any string, and those values reached code that assumes a closed set - including the prompt the AI layer builds. The check moved to the ingest door and to the single writer into the graph.
A rejected request answers 400 and names the value:
outside the ontology: node "n": unknown label "MyCustomThing"Map your events onto the labels and edge types listed in the manual. If you need a value that is not there, open an issue - adding one is a minor release, and a local string that only worked on one backend was never portable.
An apps-scoped principal can no longer act outside its applications
Section titled “An apps-scoped principal can no longer act outside its applications”Affects you if any token or OIDC claim carries an apps allowlist. A principal without
one is unaffected, and that is most deployments.
Suppressions, tickets and validations are keyed by attack-path id and were filtered by
tenant alone, while the path reads behind them were already filtered by application. So a
principal scoped to one application could suppress a path belonging to another - hiding a
real finding from the team that owns it. Those boards are now filtered, and a write against
a path outside the caller’s applications answers 404 attack path not found (or out of your scope).
If a scoped principal of yours legitimately needs a wider view, widen its apps claim; if
it needs the whole tenant, drop the claim.
One thing deliberately did not change: the tenant-wide calibration and precision/recall aggregates are still tenant-wide. They measure the engine rather than any application, name no path or asset, and GraphQL serves the same numbers - so scoping only the REST board would have been a control in name only. It is written up in the threat model.