Skip to content
kix /docs
Install the CLI

How-to guide Package task guides

Install the monitoring stack

Install Prometheus, Alertmanager, Grafana, and cluster exporters with the monitoring module.

Use kix.modules.monitoring to install the native Kix monitoring stack. The module configures Prometheus Operator, Prometheus, Alertmanager, Grafana, node-exporter, and kube-state-metrics as related package instances.

The module needs the cluster’s package catalogue. By default, buildCluster passes modules only ref and a minimal kix value. The cluster must forward the full values with specialArgs = { inherit kix packages; };, which is the first line of the snippet below. Without it the module fails to resolve the packages it instantiates.

Import the module and enable the stack in a cluster module:

how-to/package-stacks/monitoring-cluster.nix
# `kix.modules.monitoring` reads the package catalogue, which buildCluster
# does not pass to modules on its own. specialArgs forwards it.
specialArgs = { inherit kix packages; };
modules = [
kix.flavors.kind
{
imports = [ kix.modules.monitoring ];
monitoring = {
enabled = true;
prometheus = {
retention = "7d";
scrapeInterval = "30s";
};
# Keep the local example independent of persistent storage.
grafana.persistence.enabled = false;
# Add the standard Kubernetes alerts, recording rules, and monitors.
defaultRules.enabled = true;
};
# Grafana's administrator login. A real cluster keeps it in a SOPS
# source or a secret-ref instead of the cluster file.
instances.monitoring-system.grafana-admin = {
package = packages.secret;
aliases = [ "grafanaAdmin" ];
config.stringData = {
admin-user = "admin";
admin-password = "development-only-password";
};
};
}
];

View source on GitHub ↗

Grafana needs its administrator’s login: a grafanaAdmin secret in the monitoring namespace with the keys admin-user and admin-password. The example writes a development password into the cluster file. On a real cluster, keep it in a SOPS source:

cluster.nix
instances.monitoring-system.grafana-admin = {
package = packages.secret;
aliases = [ "grafanaAdmin" ];
config.source = ./secrets/grafana-admin.sops.yaml;
config.keys = [ "admin-user" "admin-password" ];
};

Without the secret, evaluation fails and names the missing grafanaAdmin dependency. The sidecars that load dashboards and datasources use the same login.

The example keeps Grafana storage ephemeral so it can run on a fresh local cluster. For a long-lived cluster, turn persistence on and size the claim:

cluster.nix
monitoring.grafana.persistence = {
enabled = true;
size = "20Gi";
};

Each persistence field has its own default (enabled = true, size = "10Gi", storageClassName = null), so setting one field keeps the others. storageClassName sets the class of Grafana’s claim; null leaves the choice to the persistent-volume-claim package.

defaultRules.enabled adds the standard Kubernetes alert rules, recording rules, and ServiceMonitors. Disable it when you want to install the core stack first and add cluster-specific rules separately.

Before deploying, adjust these settings for the cluster:

  • monitoring.prometheus.retention and retentionSize bound stored metrics.
  • instances.monitoring-system.prometheus-server.config.resources sets Prometheus CPU and memory requests and limits. Each field keeps its default until set, and null removes a default.
  • monitoring.prometheus.storage.size puts each Prometheus server’s data on a volume claim of that size, in the class of the cluster’s storageClasses provider. For another class, set instances.monitoring-system.prometheus-server.deps.storageClasses to that provider’s instance. Without a size the data lives in an emptyDir and a pod restart empties it.
  • instances.monitoring-system.prometheus-server.config.extraSpec takes any other field of the Prometheus resource, such as remoteWrite or externalUrl.
  • monitoring.alertmanager.settings is Alertmanager’s configuration file. Without it every alert goes to a receiver that drops it. The receivers’ URLs and email smarthosts decide which ports the Alertmanager pods may reach outside the cluster under network policy. The file lands in a Secret that Kix renders, so a credential written inline (auth_password, routing_key, a Slack api_url) fails evaluation: put it in a Secret, mount it with instances.monitoring-system.alertmanager.config.extraSpec.secrets, and name the file in the matching _file key.
  • monitoring.nodeExporter.enabled controls the node-exporter DaemonSet.

Grafana has no route until you set instances.monitoring-system.grafana.config.ingress.host. With that host and an ingress controller or a gateway, Grafana gets a route at the host, and that URL is its root_url. Without it, reach Grafana with kix pf. Grafana serves its own metrics at /metrics without a login, so a route makes them reachable too. To keep them private on a routed Grafana, set instances.monitoring-system.grafana.config.settings.metrics.enabled = false and instances.monitoring-system.grafana.config.metrics.enabled = false.

Set monitoring.loki.enabled to add Grafana Loki to the stack:

how-to/package-stacks/monitoring-loki-cluster.nix
monitoring = {
enabled = true;
# Keep the local example independent of a Grafana volume.
grafana.persistence.enabled = false;
# Loki runs beside Grafana on a claim named loki-storage. Grafana
# gets a Loki datasource, and the module's log shipper sends every
# pod's logs to it.
loki = {
enabled = true;
storageSize = "5Gi";
};
};
# Loki's own configuration, merged over the package's.
instances.${config.monitoring.namespace}.loki.config.settings = {
limits_config.retention_period = "72h";
};

View source on GitHub ↗

The module runs Loki as one pod in the monitoring namespace, on a PersistentVolumeClaim instance named loki-storage of monitoring.loki.storageSize (10Gi by default). Size it for the log volume you keep: Loki holds every line for the retention period, 7 days unless settings.limits_config.retention_period says otherwise. To pick a StorageClass, set instances.<monitoring namespace>.loki-storage.config.storageClassName.

Loki’s own configuration goes in instances.<monitoring namespace>.loki.config.settings, in Loki’s config.yaml keys. A few keys the package needs (the listen port, the data paths, and single tenancy) are locked, and setting them fails evaluation with the name of the key. The same goes for the extraArgs flags that would override those keys. Grafana gets a Loki datasource from the Loki instance, with a derived network policy that lets it query Loki, and monitoring.lokiLogs.enabled (on with Loki) runs the loki-logs agent on every node, at the module’s log-shipper PriorityClass, to send every pod’s logs to it.

The agent reads the node’s log files as root through host paths, and node-exporter mounts host paths too, so the monitoring namespace runs at the privileged Pod Security level. For a restricted monitoring namespace, turn off monitoring.nodeExporter.enabled and monitoring.lokiLogs.enabled and declare those agents in a namespace of their own.

The loki package takes its volume through a storage dependency. Declare a persistent-volume-claim instance with a size beside it; the header of kix/packages/loki/default.nix shows the form. When another package in the same namespace has an optional storage dependency (Grafana does), give Loki’s claim its own instance name and aliases, and wire it with deps.storage = ref.<namespace>.<claim>. Otherwise both pods mount the one ReadWriteOnce claim.

Evaluate all monitoring resources:

kix-examples/
❱ kix check how-to-package-monitoring
 TOOL       RESULT  DETAILS                      
 eval       pass    110 manifests evaluated      
 scorecard  pass    0 errors, 6 warnings, 4 info

List the package instances selected by the module:

kix-examples/
❱ kix list packages --cluster how-to-package-monitoring
 NAME                 VERSION  OWNER     STATUS                        
 alertmanager         v0.33.1  platform  installed [monitoring-system] 
 default-rules        -        platform  installed [monitoring-system] 
 grafana              13.1.1   platform  installed [monitoring-system] 
 grafana-admin        1.0.0    platform  installed [monitoring-system] 
 kube-state-metrics   v2.19.1  platform  installed [monitoring-system] 
 node-exporter        v1.12.1  platform  installed [monitoring-system] 
 platform-dns         -        -         import [kube-system]          
 platform-storage     -        -         import [kube-system]          
 prometheus-operator  v0.92.1  platform  installed [monitoring-system] 
 prometheus-server    v3.13.1  platform  installed [monitoring-system]

The list includes the operator and the components it manages. The flavor’s DNS and storage entries are imports supplied by the target platform.

Deploy the cluster, then check readiness:

kix-examples/
❱ kix deploy how-to-package-monitoring
❱ kix status how-to-package-monitoring

On a small local cluster, Prometheus and Grafana can take several minutes to become ready while their images are pulled.

To open Grafana locally, forward its port:

kix-examples/
❱ kix pf how-to-package-monitoring grafana 3000:3000

Open http://127.0.0.1:3000 while the forward is running and sign in with the login from the grafanaAdmin secret. Stop the forward with Ctrl-C.

If a component remains pending, use kix status to find its workload, then check node capacity, PVC provisioning, and any pod security restrictions on the target cluster.