Skip to content

Release notes ClusterControl MCP

The ClusterControl MCP Server is versioned and released independently of the main ClusterControl components, as clustercontrol-mcp-1.0.0-N builds. This page is the running release history for the MCP Server. The initial launch (clustercontrol-mcp-1.0.0-5, May 4th, 2026) is documented in the v2.4.0 release notes.

Maintenance Release: September 30th, 2026

  • Build:
    • clustercontrol-mcp-1.0.0-124

This release fixes how the server reads backup records, so the backup tools list real backups again. The tools, their arguments and the server's settings are unchanged from 1.0.0-122.

  • Backups are listed with their real id, method, status, timestamps and size again. list_backups, get_latest_backup and the clustercontrol://clusters/{cluster_id}/backups resource showed a backup that exists as #0 with a size of 0 B and no method, status or timestamps, so a restore or a delete could not be pointed at a real backup id. The server asks the controller for version-2 backup records, which the controller returns nested under metadata, but it read them in the older flat shape. It now reads both shapes, and the output format is unchanged. Every earlier build was affected, against ClusterControl 2.4.0 and 2.5.0 controllers alike. (CLUS-8812)

Feature Release: Advisors that report findings, and Terraform export

  • Build:
    • clustercontrol-mcp-1.0.0-122

Where the previous release made the server safe to operate, this one lets it advise. Six new tools inspect a cluster and report what is wrong or risky — as machine-readable findings a client can act on, not prose — and an existing cluster can now be exported as Terraform. Every new tool is read-only, so read-only mode, tool scoping and the audit log apply to them exactly as to any other read tool.

⚠ Read before upgrading

No tool, argument or setting was removed or renamed. Three changes can still affect a client that parses output strictly; all three are described under Behaviour changes below.

  • lag=0s becomes lag=unknown in topology and incident-bundle text wherever the controller reported no lag.
  • Cluster host JSON gains keys.
  • With MCP_TOOL_ALLOW set, the six new tools are out of scope until you add them.
  • One findings contract for every advisor — each advisor returns {findings, summary, catalog_version}, both as structuredContent with a published outputSchema and as the same JSON in the text content, so a client can key on a finding's code instead of parsing prose. Codes, severities (info, warn, critical), recommendations and thresholds come from a versioned catalog that ships inside the binary (version 7), so changing a threshold is a catalog version bump rather than a silent change in behaviour. A healthy cluster produces no warn or critical findings. (CLUS-8116)
  • advise_replication_topology — broken replication links (REPL_LINK_DOWN), replicas more than 60 seconds behind (REPL_LAG_HIGH), and a cluster with a single database node (SINGLE_NODE_PRIMARY). MySQL replication and PostgreSQL streaming replication are each judged by their own model, and an unknown lag is reported as unknown, never as zero. (CLUS-8113)
  • audit_db_privileges — accounts holding every global privilege (PRIV_GLOBAL_ALL), accounts that accept connections from any host (PRIV_WILDCARD_HOST), and more than three superuser accounts (PRIV_SUPERUSER_COUNT_HIGH), on MySQL-family and PostgreSQL clusters. System accounts are excluded. (CLUS-8116)
  • check_config_drift — a configuration parameter whose value differs between nodes of the same role (CONFIG_DRIFT), with each node's value and the node that is off the majority. Parameters that differ per node by design — server_id, addresses, paths — are never compared. (CLUS-8117)
  • check_ssl_posture — SSL turned off on a database host (SSL_DISABLED), and the state of each of the cluster's certificates from ClusterControl's own CA: revoked, expired, expiring within 30 days, valid, or unknown. On MySQL-family hosts it asks the database server itself whether TLS is available rather than trusting the controller's flag alone — see Bug fixes. (CLUS-8115, CLUS-8762)
  • check_version_eol — whether each database release is supported, ending within 90 days, past community maintenance, or end of life, resolved per engine against your support entitlement. The lifecycle table covers 30 releases of MySQL, MariaDB and MongoDB, ships inside the package and is never fetched, so the check works in an air-gapped installation; every result carries the table's date so its age is visible. (CLUS-8118, CLUS-8767)
  • Support entitlement settings (MCP_SUPPORT_ENTITLEMENT, MCP_EOL_TABLE_FILE) — read only by check_version_eol, and both unset by default. The entitlement is given per engine and optionally per cluster, for example mysql=premier,mariadb=community,cluster:7=extended; unset means each engine's free community stream. The table file replaces the bundled lifecycle table, for testing or to carry newer dates before the next release. Both fail fast: a value the engine does not accept, or a table file that cannot be read, stops the server at startup with an error naming the entry. (CLUS-8118)
  • export_terraform — an existing cluster as Terraform HCL for the severalnines/clustercontrol provider v0.2.25, returned in the tool result; the server writes no file. Credentials appear only as var.* variables. MySQL replication, Galera and PostgreSQL replication clusters are exported; an engine or layout the provider cannot express is refused with EXPORT_UNSUPPORTED_ENGINE or EXPORT_UNSUPPORTED_TOPOLOGY rather than exported as a different cluster. (CLUS-8119)
  • The documentation bundle that search_docs serves offline is refreshed from the current ClusterControl documentation. (CLUS-8769)
  • The tool registry grows from 74 to 80 tools (51 pure-read, 26 CMON-mutating, 3 local-filesystem-writing; 22 flagged destructive), adding the six read-only tools above. The 21 resource surfaces are unchanged.
  • lag=0s becomes lag=unknown in get_cluster_topology and get_incident_bundle text wherever the controller reported no lag, and the topology view gains a lag_unknown field. Anything that read lag=0s as "caught up" was being told something the controller never said. (CLUS-8113)
  • Cluster host JSON gains keys where the controller reports them: version, hostname_data, hostname_internal, ip, synchronous, sync_state, and the replication link status with its IO/SQL error numbers and states. A client that rejects unknown keys needs to accept them. (CLUS-8113, CLUS-8119)
  • With MCP_TOOL_ALLOW set, the new tools are out of scope until they are added to the list. That is the scope working as designed.
  • The configuration readers hide more parameters — every parameter whose name contains conninfo is now hidden. See Security fixes. (CLUS-8766)
  • The configuration readers no longer return replication credentials in clear text. get_node_config, get_cluster_config and the configuration resource returned the wsrep-provider-options and loose-wsrep-provider-options spellings, and PostgreSQL primary_conninfo, unmasked — values that can carry a replication password. Parameter names are now normalised before the masking decision, and every name containing conninfo is hidden, which deliberately also hides some non-secret parameters. (CLUS-8766)
  • set_node_config works on MySQL-family clusters. It sent no configuration group, so outside CCX mode the controller rejected the change with "Group is missing.", and unset_node_config named no section of my.cnf. Both tools now take an optional group; when it is omitted, a MySQL, Galera or Group Replication node gets mysqld and any other node gets none. (CLUS-8760)
  • check_ssl_posture no longer reports SSL enabled on a MySQL-family server that refuses TLS. The controller derives its SSL flag from tls_version, which lists protocols even when the server cannot negotiate TLS; this was reproduced on MySQL 8.0, MySQL 8.4 and MariaDB 11.4. Where the controller reports a MySQL-family host enabled, the tool now reads the server's own have_ssl, or on MySQL 8.4 whether a certificate is loaded, and reports SSL_DISABLED when the server says TLS is off. The check only ever turns enabled into disabled; a host that does not answer keeps the controller's flag and is named in the summary. (CLUS-8762)
  • Reading a PostgreSQL cluster no longer fails on an unusual lag value. A replication lag reported as "--" or as a fraction broke the whole cluster read for every tool, and an empty lag no longer reads as 0. (CLUS-8113)
  • The advisors read the controller's sampled state. The replication advisor sees the controller's last sample, taken about every 10 seconds by default, and REPL_LAG_HIGH uses its own 60-second threshold, separate from the controller's max_replication_lag alarm. Galera is judged for redundancy only: REPL_LINK_DOWN and REPL_LAG_HIGH are never raised on a Galera cluster, and the summary says so.
  • Two replication codes are reserved but never emitted. NO_FAILOVER_CANDIDATE and REPL_MONITOR_GRANT_MISSING are in the catalog, but the controller exposes no clean source for either, so promotability and the monitoring account's grants are not assessed.
  • check_config_drift compares configuration files, not running servers. A parameter changed with SET GLOBAL, or a file edited but not reloaded, is not seen as the server has it, and memory-sized settings legitimately differ between nodes on different hardware. Hidden parameters are never compared.
  • check_version_eol assesses database nodes, never the controller, and covers 30 releases. PostgreSQL, Percona, Redis/Valkey, SQL Server, and releases the table does not list return EOL_DATA_UNAVAILABLE. The table's dates change only with a new package. Oracle Sustaining Support does not count as supported, so MySQL 8.0 reads VERSION_EOL even under mysql=extended.
  • audit_db_privileges is literal. A stock MySQL cluster always reports PRIV_GLOBAL_ALL, and because accounts are counted per user@host it can exceed the superuser limit. A single-node MySQL cluster is refused, because the controller does not list its accounts; PostgreSQL host rules are not assessed for PRIV_WILDCARD_HOST; and privileges held only through a granted role, and dynamic privileges, are not seen.
  • check_ssl_posture reads certificate expiry from ClusterControl's CA, so a certificate it did not issue or import is reported SSL_CERT_EXPIRY_UNKNOWN.
  • export_terraform describes how to recreate a cluster, not how to manage it. The provider cannot import an existing cluster, so terraform plan against an export is never a no-op; applying it against a different controller deploys onto the same hosts, even while they run this cluster; and the Terraform state holds the passwords you supply in plaintext. ClickHouse clusters and load balancers are not exported. No export from this release has been applied against a controller.
  • ClickHouse is not covered by the advisors in this release.
  • Packages are x86_64 only in this release.

Feature Release: Per-session identity, a confirmation step, and job forensics

  • Build:
    • clustercontrol-mcp-1.0.0-108

Where the previous release let the server refuse actions, this one lets it tell two callers apart. Each assistant can now hold its own CMON identity, a destructive call can be made to require a confirmation — or a person — before it runs, and a failed job can be asked what it was doing when it stopped.

⚠ Breaking changes

Three changes alter behaviour a working deployment may depend on. All three are described under Behaviour changes below.

  • On the HTTP transport, session IDs the server did not issue are now refused, and outstanding IDs do not survive a restart. Two instances behind a load balancer need sticky routing.
  • create_job now refuses a command that is not in this build's catalogue.
  • get_node_config and get_controller_config hide more parameters than before.
  • Per-session identity (MCP_IDENTITY_MAP_FILE) — an operator-owned map binds each bearer token to its own CMON username and credential, so the controller's audit log attributes each assistant's actions to a different user. A token that is not in the map resolves to nothing: there is no default entry and no fall back to the shared credentials, so an unmapped token is refused rather than arriving as somebody else. A map that does not parse, or that is ambiguous, is a startup error rather than a server that silently dropped an entry. The map requires the HTTP transport — identity rides on the Authorization header, and selecting a map on stdio is a startup refusal. (CLUS-8102, CLUS-8646)
  • get_session_info now reports whether that identity actually works. It returns one of no_identity_bound, not_attempted, verified or failed, plus identity_last_error carrying the controller's own sentence when it refused. Previously the tool reported the configured name and nothing else, so a map naming a CMON user that does not exist looked healthy until the first call failed. Credentials are still not probed at startup — the verdict is what CMON answered on a call the server has already made. (CLUS-8738)
  • Two-step execution guard (MCP_EXECUTION_GUARD=enforced) — a destructive call is admitted only against a confirmation_token the server minted on a preview of the same operation. The confirmation covers one identity, one session, one tool and one set of arguments; it is single-use, it expires, and a restart revokes every outstanding one. A preview is never refused, and is where the confirmation comes from. (CLUS-8108)
  • Human-approval tier (MCP_HUMAN_APPROVAL=elicitation) — above the guard, and usable on its own. A destructive execution is put to a person over the calling session's elicitation channel, carrying the operation's digest and a single-use code the answer must return. Where an approval cannot be sought at all — no bound identity, or a session that cannot be elicited — the call is refused with HUMAN_APPROVAL_UNAVAILABLE rather than proceeding or hanging. (CLUS-8108)
  • Standard operating procedure prompts — sop_failover_drill and sop_major_upgrade, served as MCP prompts. Each previews before it acts and carries an explicit branch for a failure part-way through. sop_major_upgrade opens with an engine guard (PostgreSQL only) and a backup check; sop_failover_drill opens with a preflight. Both name a human-approval step, which this server enforces only when MCP_HUMAN_APPROVAL is on — it says so at startup when it is not. (CLUS-8109, CLUS-8110)
  • get_job_checkpoint — where one job actually got to: its command, status, whether this build recognises that status at all, and the tail of its log. Fetched by job ID alone, so a job older than one page of list_jobs is still answered for. It never suggests re-submitting a job: CMON jobs are not idempotent, so every hint it gives is a query or a wait. (CLUS-8111)
  • get_incident_bundle can be anchored on a job. Passing job_id adds a failed_job section and widens the jobs window to reach that job — never narrowing it below what you asked for. Only the jobs section is job-scoped, and the other sections say why they cannot be. (CLUS-8112)
  • Client capability probe — the server records what each client declared at connection time and consults it when a control needs to reach a person, so an approval-required call is refused on a session that cannot be elicited rather than waiting for an answer that cannot arrive. (CLUS-8107)
  • The tool registry grows from 72 to 74 tools (45 pure-read, 26 CMON-mutating, 3 local-filesystem-writing; 22 flagged destructive), adding get_session_info and get_job_checkpoint. The 21 resource surfaces are unchanged.
  • Session IDs are validated for issuance. On the HTTP transport the server now admits only the Mcp-Session-Id values it issued, where previously any correctly shaped value was served. Two consequences: outstanding IDs do not survive a restart, so clients get 404 and must re-initialize; and two server instances require sticky routing for a session's full lifetime, because a session established against one is not known to the other. (CLUS-8646)
  • With an identity map configured, /sse and /message are not served. SSE's separate session mechanism sits outside the identity binding, so it is refused rather than left as a way around it. Deployments using SSE should stay on the shared-token configuration until their clients speak the streamable HTTP transport. (CLUS-8102)
  • create_job refuses a command outside this build's catalogue. Previously such a call was accepted and reported a job ID — CMON creates a job titled Unknown Command, runs nothing and fails it — so a caller was told work had started that never would. The catalogue is a snapshot taken from the controller source at build time, so a command added by a newer controller is also refused; the refusal says the command is not in this build's catalogue, and never that it is not a real CMON command. (CLUS-8736)
  • get_node_config and get_controller_config hide more parameters. The rule that decides what counts as a credential is now shared with the approval prompt and matches key and auth as substrings, so operational parameters whose names contain them — key_buffer_size, foreign_key_checks, authentication_policy and similar — are now returned as <hidden>. This over-matching is deliberate: the alternative is a hand-maintained exemption list whose failure mode is disclosure. (CLUS-8740)
  • Task status notifications no longer reach every connected client. A background job's progress was broadcast to all sessions, so one assistant could observe another's activity. Notifications are now scoped per session when an identity map is configured; without a map every caller is the same CMON principal, which is the pre-existing behaviour. (CLUS-8737)
  • get_node_config no longer returns credentials verbatim. wsrep_sst_auth carries sstuser:<password> and matched none of the previous rules, so Galera SST credentials were returned in clear text. The read path is now no wider than the approval prompt, from one shared list. (CLUS-8740)
  • The destructive-operation warning now fires for the commands it names. Four of the nine job commands it escalated on were names CMON does not accept — shutdown, remove_node, drop_cluster and rebuild_replication_slave — so the extra warning could never appear for stopping or removing a cluster. The same fabricated names appeared in the operator-facing scope warning. Both are corrected against the controller's own command list. (CLUS-8736, CLUS-8741)
  • A tool declaring that it writes to the filesystem can no longer bypass the export directory. The build-time check enforced only that a tool declaring no local writes made none; a tool declaring the opposite could call the filesystem directly. Both directions are now enforced. (CLUS-8707)
  • A non-credential refusal is recorded as a credential failure. identity_verification reads failed whenever the controller answers AccessDenied, and CMON answers AccessDenied both for a wrong password and for a licensing refusal or a controller whose database is unavailable. An identity whose credentials are correct can therefore read failed while the controller is starting up. (CLUS-8752)
  • Scoping removes a tool, not a data class. Several tools read the same data, so denying one can leave a sibling serving the same rows. The server names each case at startup as tool scope: WARNING: ….
  • create_job still reaches job commands no tool exposes. It is now bounded — the command must be in the committed catalogue — but that catalogue is wider than the tool surface and includes destructive commands such as remove_cluster, removenode and failover. The startup warning about this is narrower than the exposure: it appears only when a tool scope is configured, read-only mode is off, create_job is in scope, and that scope denies a job-backed tool. A deployment with no scope configured has the exposure in full and gets no warning.
  • Symlink protection covers the final path component only. For the digest key and the audit log, a hard link at the path itself, or a symlink at any parent directory, is followed normally. Whoever can write a directory on those paths controls what is at the end of them.
  • Packages are x86_64 only in this release.

Feature Release: Access controls, audit logging and offline docs

  • Build:
    • clustercontrol-mcp-1.0.0-90

This release adds the controls needed to connect an AI assistant to a production controller: the server can now refuse actions itself, rather than relying on the client to respect an advisory hint.

  • Global read-only mode (MCP_READ_ONLY) — refuses all 29 state-changing tools for the lifetime of the process. Enforced on two independent layers: the write tools are hidden from the client's tool list, and a direct call on one is refused, so a client that calls a tool it was never shown is still refused. (CLUS-8099)
  • Per-tool and per-resource scoping (MCP_TOOL_ALLOW, MCP_TOOL_DENY) — expose only the tools you name. Deny wins over allow, the two compose with read-only mode, and resource surfaces follow their tools automatically. An entry matching no tool is a startup error rather than a silent no-op. (CLUS-8100)
  • Audit log (MCP_AUDIT_LOG) — every tool call and resource read recorded twice as JSONL, an intent record before the action and a completion after, with a third acceptance record on the background-task path. Raw arguments are never written; what is recorded is a keyed digest, so a reader of the log cannot work backwards to a password. Auditing is fail-closed: if the intent record cannot be written, the action does not run. (CLUS-8103)
  • Export directory containment (MCP_EXPORT_DIR) — the three tools that write to the MCP Server's own filesystem are confined to one directory, whatever path the client asks for. (CLUS-8581)
  • search_docs — ranked excerpts from a ClusterControl documentation bundle shipped inside the package. No network calls, so it works in an air-gapped installation. (CLUS-8104, CLUS-8105)
  • get_incident_bundle — one read-only call returning alarms, recent jobs, collected log tails, a metric summary with outlier flags, and topology. Every section is always present and explicitly marked [error], [empty] or [truncated], so a section missing because a source failed cannot be mistaken for one that was empty. (CLUS-8106)
  • The tool registry grows from 70 to 72 tools (43 pure-read, 26 CMON-mutating, 3 local-filesystem-writing; 22 flagged destructive), with 21 resource surfaces unchanged.
  • Local file writes are confined by default. A server whose configuration is untouched will begin refusing destination paths outside /var/lib/cmon-mcp/exports — including paths such as /tmp that previously worked. Set MCP_EXPORT_DIR to the directory your callers already use to keep their existing paths valid. (CLUS-8581)
  • A fresh install is audited by default. First installs provision a digest key and switch auditing on. Upgrades enable nothing, because a host that has never been audited has no key and a server told to audit without one refuses to start. Set MCP_AUDIT_LOG=off in /etc/default/cmon-mcp to opt out. (CLUS-8583)
  • The configuration file is /etc/default/cmon-mcp. It follows the service and binary names; the package is still clustercontrol-mcp. On RPM hosts an earlier build could leave settings behind in a .rpmsave file during this rename — the installer now reports exactly what it found and where. (CLUS-8601, CLUS-8648)
  • A flag-parse error no longer reprints the CMON password and MCP bearer token to the journal. (CLUS-8587)
  • /etc/default/cmon-mcp, which holds CMON_PASSWORD and MCP_AUTH_TOKEN, now ships mode 0600 rather than world-readable, and the installer tightens every copy of it the package manager leaves behind. (CLUS-8594)
  • export_cluster_log is no longer an arbitrary-file write as root; every client-supplied path is confined to the export directory and lands through a staged rename, so a symlink at the destination is never written through. (CLUS-8581)
  • get_controller_config no longer returns live Vault and OpenBao tokens in clear text. (CLUS-8650)
  • Credentials and key material are held where Go's formatting verbs cannot reach them, so no log line or error message can print them by accident. (CLUS-8595)
  • The digest key file is no longer symlink-swappable, and governed mode refuses to write its marker through a dangling symlink. (CLUS-8576, CLUS-8647)
  • Every session runs as one CMON identity. The server authenticates as the single configured CMON_USERNAME, and the audit log attributes every action to it, so two assistants cannot be told apart from the record. Per-session identity is planned for a later release.
  • Scoping removes a tool, not a data class. Several tools read the same data, so denying one can leave a sibling serving the same rows. The server names each case at startup as tool scope: WARNING: ….
  • create_job reaches the whole job surface. Its free-form command argument reaches CMON job commands that have no tool of their own and that no deny list can name.
  • Packages are x86_64 only in this release.

Maintenance Release: September 1st, 2026

  • Build:
    • clustercontrol-mcp-1.0.0-60

⚠ Breaking change

get_cluster_log no longer accepts output_path and no longer writes files — it is now a pure read that returns log content (the tail of the last 200 lines by default). Callers that used output_path to save a log file locally must switch to the new export_cluster_log tool, which requires output_path and writes the file with mode 0600. (CLUS-8098)

  • export_cluster_log — exports a collected cluster log file in full to a local output_path on the MCP Server host. The file is written with mode 0600, including when overwriting an existing file that had looser permissions. (CLUS-8098)
  • Tool annotations — all 70 tools now serve the standard MCP readOnlyHint and destructiveHint annotations. The annotations are generated at tool registration from a committed, human-reviewed manifest, so the served hints cannot drift from the manifest. Note that annotations are advisory metadata for MCP clients — clients can use them to auto-approve reads and require confirmation before destructive calls — they are not server-side enforcement, and dry_run previews remain a client-side safeguard. (CLUS-8098)
  • The tool registry grows from 69 to 70 tools (41 pure-read, 26 CMON-mutating, 3 local-filesystem-writing; 22 flagged destructive), with 21 resource surfaces (3 static + 18 templates). (CLUS-8098)