How-to guide Package task guides
Install the monitoring stack
Install Prometheus, Alertmanager, Grafana, and cluster exporters with the monitoring module.
Use kix.modules.monitoring to install the native Kix monitoring stack. The
module configures Prometheus Operator, Prometheus, Alertmanager, Grafana,
node-exporter, and kube-state-metrics as related package instances.
The module needs the cluster’s package catalogue. By default, buildCluster
passes modules only ref and a minimal kix value. The cluster must forward
the full values with specialArgs = { inherit kix packages; };, which is the
first line of the snippet below. Without it the module fails to resolve the
packages it instantiates.
Add the monitoring module
Section titled “Add the monitoring module”Import the module and enable the stack in a cluster module:
# `kix.modules.monitoring` reads the package catalogue, which buildCluster # does not pass to modules on its own. specialArgs forwards it. specialArgs = { inherit kix packages; };
modules = [ kix.flavors.kind { imports = [ kix.modules.monitoring ];
monitoring = { enabled = true; prometheus = { retention = "7d"; scrapeInterval = "30s"; };
# Keep the local example independent of persistent storage. grafana.persistence.enabled = false;
# Add the standard Kubernetes alerts, recording rules, and monitors. defaultRules.enabled = true; }; # Grafana's administrator login. A real cluster keeps it in a SOPS # source or a secret-ref instead of the cluster file. instances.monitoring-system.grafana-admin = { package = packages.secret; aliases = [ "grafanaAdmin" ]; config.stringData = { admin-user = "admin"; admin-password = "development-only-password"; }; }; } ];Grafana needs its administrator’s login: a grafanaAdmin secret in the
monitoring namespace with the keys admin-user and admin-password. The
example writes a development password into the cluster file. On a real
cluster, keep it in a SOPS source:
instances.monitoring-system.grafana-admin = { package = packages.secret; aliases = [ "grafanaAdmin" ]; config.source = ./secrets/grafana-admin.sops.yaml; config.keys = [ "admin-user" "admin-password" ];};Without the secret, evaluation fails and names the missing grafanaAdmin
dependency. The sidecars that load dashboards and datasources use the same
login.
The example keeps Grafana storage ephemeral so it can run on a fresh local cluster. For a long-lived cluster, turn persistence on and size the claim:
monitoring.grafana.persistence = { enabled = true; size = "20Gi";};Each persistence field has its own default (enabled = true,
size = "10Gi", storageClassName = null), so setting one field keeps the
others. storageClassName sets the class of Grafana’s claim; null leaves the
choice to the persistent-volume-claim package.
defaultRules.enabled adds the standard Kubernetes alert rules, recording
rules, and ServiceMonitors. Disable it when you want to install the core stack
first and add cluster-specific rules separately.
Before deploying, adjust these settings for the cluster:
monitoring.prometheus.retentionandretentionSizebound stored metrics.instances.monitoring-system.prometheus-server.config.resourcessets Prometheus CPU and memory requests and limits. Each field keeps its default until set, and null removes a default.monitoring.prometheus.storage.sizeputs each Prometheus server’s data on a volume claim of that size, in the class of the cluster’sstorageClassesprovider. For another class, setinstances.monitoring-system.prometheus-server.deps.storageClassesto that provider’s instance. Without a size the data lives in an emptyDir and a pod restart empties it.instances.monitoring-system.prometheus-server.config.extraSpectakes any other field of the Prometheus resource, such asremoteWriteorexternalUrl.monitoring.alertmanager.settingsis Alertmanager’s configuration file. Without it every alert goes to a receiver that drops it. The receivers’ URLs and email smarthosts decide which ports the Alertmanager pods may reach outside the cluster under network policy. The file lands in a Secret that Kix renders, so a credential written inline (auth_password,routing_key, a Slackapi_url) fails evaluation: put it in a Secret, mount it withinstances.monitoring-system.alertmanager.config.extraSpec.secrets, and name the file in the matching_filekey.monitoring.nodeExporter.enabledcontrols the node-exporter DaemonSet.
Grafana has no route until you set
instances.monitoring-system.grafana.config.ingress.host. With that host and
an ingress controller or a gateway, Grafana gets a route at the host, and
that URL is its root_url. Without it, reach Grafana with kix pf. Grafana
serves its own metrics at /metrics without a login, so a route makes them
reachable too. To keep them private on a routed Grafana, set
instances.monitoring-system.grafana.config.settings.metrics.enabled = false
and instances.monitoring-system.grafana.config.metrics.enabled = false.
Add Loki
Section titled “Add Loki”Set monitoring.loki.enabled to add Grafana Loki to the stack:
monitoring = { enabled = true;
# Keep the local example independent of a Grafana volume. grafana.persistence.enabled = false;
# Loki runs beside Grafana on a claim named loki-storage. Grafana # gets a Loki datasource, and the module's log shipper sends every # pod's logs to it. loki = { enabled = true; storageSize = "5Gi"; }; };
# Loki's own configuration, merged over the package's. instances.${config.monitoring.namespace}.loki.config.settings = { limits_config.retention_period = "72h"; };The module runs Loki as one pod in the monitoring namespace, on a
PersistentVolumeClaim instance named loki-storage of
monitoring.loki.storageSize (10Gi by default). Size it for the log volume
you keep: Loki holds every line for the retention period, 7 days unless
settings.limits_config.retention_period says otherwise. To pick a
StorageClass, set
instances.<monitoring namespace>.loki-storage.config.storageClassName.
Loki’s own configuration goes in
instances.<monitoring namespace>.loki.config.settings, in Loki’s
config.yaml keys. A few keys the package needs (the listen port, the data
paths, and single tenancy) are locked, and setting them fails evaluation
with the name of the key. The same goes for the extraArgs flags that would
override those keys. Grafana gets a Loki datasource from the Loki instance,
with a derived network policy that lets it query Loki, and
monitoring.lokiLogs.enabled (on with Loki) runs the loki-logs agent on
every node, at the module’s log-shipper PriorityClass, to send every pod’s
logs to it.
The agent reads the node’s log files as root through host paths, and
node-exporter mounts host paths too, so the monitoring namespace runs at the
privileged Pod Security level. For a restricted monitoring namespace,
turn off monitoring.nodeExporter.enabled and monitoring.lokiLogs.enabled
and declare those agents in a namespace of their own.
Without the module
Section titled “Without the module”The loki package takes its volume through a storage dependency. Declare a
persistent-volume-claim instance with a size beside it; the header of
kix/packages/loki/default.nix shows the form. When another package in the
same namespace has an optional storage dependency (Grafana does), give
Loki’s claim its own instance name and aliases, and wire it with
deps.storage = ref.<namespace>.<claim>. Otherwise both pods mount the one
ReadWriteOnce claim.
Check the stack
Section titled “Check the stack”Evaluate all monitoring resources:
❱ kix check how-to-package-monitoring
TOOL RESULT DETAILS
eval pass 110 manifests evaluated
scorecard pass 0 errors, 6 warnings, 4 info List the package instances selected by the module:
❱ kix list packages --cluster how-to-package-monitoring
NAME VERSION OWNER STATUS
alertmanager v0.33.1 platform installed [monitoring-system]
default-rules - platform installed [monitoring-system]
grafana 13.1.1 platform installed [monitoring-system]
grafana-admin 1.0.0 platform installed [monitoring-system]
kube-state-metrics v2.19.1 platform installed [monitoring-system]
node-exporter v1.12.1 platform installed [monitoring-system]
platform-dns - - import [kube-system]
platform-storage - - import [kube-system]
prometheus-operator v0.92.1 platform installed [monitoring-system]
prometheus-server v3.13.1 platform installed [monitoring-system] The list includes the operator and the components it manages. The flavor’s DNS and storage entries are imports supplied by the target platform.
Deploy and inspect
Section titled “Deploy and inspect”Deploy the cluster, then check readiness:
❱ kix deploy how-to-package-monitoring❱ kix status how-to-package-monitoring On a small local cluster, Prometheus and Grafana can take several minutes to become ready while their images are pulled.
To open Grafana locally, forward its port:
❱ kix pf how-to-package-monitoring grafana 3000:3000 Open http://127.0.0.1:3000 while the forward is running and sign in with
the login from the grafanaAdmin secret. Stop the forward with Ctrl-C.
If a component remains pending, use kix status to find its workload, then
check node capacity, PVC provisioning, and any pod security restrictions on
the target cluster.