example(ske): add example for kubeapi audit logs #65

Merged
mauritz.uphoff merged 4 commits from example/ske-kubeapi-audit-log into main 2026-09-22 08:24:04 +00:00
Owner

Hey @mauritz.uphoff ,

I could not test this today end-to-end as we have issues with stackit_telemetrylink (see https://chat.google.com/room/AAQAyBUdV7k/V2pak9TVhOY/V2pak9TVhOY?cls=10)

But I almost rolled out similar but with STACKIT Logs and without automated k8s ressource deployments.

Main "issue" is that we can not filter (irrelevant) logs (see readme) on Telemetry Router level. Filter on SKE side is also not possible as discussed with the team.

Hey @mauritz.uphoff , I could not test this today end-to-end as we have issues with stackit_telemetrylink (see https://chat.google.com/room/AAQAyBUdV7k/V2pak9TVhOY/V2pak9TVhOY?cls=10) But I almost rolled out similar but with STACKIT Logs and without automated k8s ressource deployments. Main "issue" is that we can not filter (irrelevant) logs (see readme) on Telemetry Router level. Filter on SKE side is also not possible as discussed with the team.
example(ske): add example for kubeapi audit logs
Some checks failed
Default CI / Check for Open TODOs (pull_request) Successful in 1m40s
AI PR Review / AI PR Review (pull_request) Successful in 3m29s
Default CI / Pre-Commit Hooks (pull_request) Failing after 11m35s
75c5951542

🤖 AI PR Review

75c59515 · STACKIT Model Serving & STACKIT Cloud Advisor

⚠️ 📝 Spelling & Grammar
🤖 STACKIT Model Serving
  • examples/ske-kubeapi-audit-log/README.md:15 — “comming” → “coming”
  • examples/ske-kubeapi-audit-log/README.md:15 — “overwhleming” → “overwhelming”
  • examples/ske-kubeapi-audit-log/README.md:24 — “review” → “preview” (in “private review” → should be “private preview” to match context)
  • examples/ske-kubeapi-audit-log/README.md:59 — “STACKIT Logs only focusing on have a managed Loki” → “STACKIT Logs focuses only on providing a managed Loki”
  • examples/ske-kubeapi-audit-log/README.md:60 — “Observability comes with a managed Grafana on top.” → “Observability includes a managed Grafana on top.” (for smoother phrasing)
  • examples/ske-kubeapi-audit-log/README.md:81 — “The resource attributes STACKIT sets on every forwarded record (dots become underscores once they arrive as Loki labels):” → “The resource attributes STACKIT sets on every forwarded record (dots become underscores when they arrive as Loki labels):” (more natural phrasing)
🔍 STACKIT Cloud Advisor

Spelling and Grammar Review

I have reviewed the provided git diff for spelling and grammar errors in the prose content, including Markdown documentation, Terraform variable descriptions, and inline comments. I have also verified the usage of STACKIT product names and technical terminology against the provided documentation.

The following errors were identified:

File Line Wrong Text Correction
examples/ske-kubeapi-audit-log/010-provider.tf 18 distributed under an "AS IS" BASIS distributed on an "AS IS" BASIS
examples/ske-kubeapi-audit-log/020-variables.tf 64 Access control list of the Observability/Logs. Access control list of the Observability/Logs. (Note: While grammatically acceptable, "for" is more idiomatic: Access control list for Observability/Logs.)
examples/ske-kubeapi-audit-log/080-canary-workload.tf 13 into the audit into the audit stream (Sentence is truncated/incomplete)
examples/ske-kubeapi-audit-log/README.md 11 comming from kublet or gardener coming from kubelet or gardener
examples/ske-kubeapi-audit-log/README.md 11 This can be overwhleming This can be overwhelming
examples/ske-kubeapi-audit-log/README.md 15 in private review in private preview
examples/ske-kubeapi-audit-log/README.md 53 STACKIT Logs only focusing on have a managed Loki STACKIT Logs only focus on having a managed Loki
examples/ske-kubeapi-audit-log/README.md 108 where the kublet/gardener noise where the kubelet/gardener noise
⚠️ 🏗️ Infrastructure Changes
🤖 STACKIT Model Serving
  • Creates STACKIT network for SKE nodes with IPv4 prefix and nameservers
  • Creates STACKIT Observability instance (Loki+Grafana) with log retention and ACL
  • Creates STACKIT Telemetry Router instance with access token and destination config
  • Creates STACKIT Telemetry Link to route project audit logs to the router
  • Creates SKE Kubernetes cluster with audit enabled, using Flatcar OS and node pools
  • Creates ephemeral kubeconfig for cluster access during Terraform apply
  • Deploys Kubernetes namespace and ConfigMap as audit canary workload via ephemeral kubeconfig
  • Exposes outputs for cluster name, router IDs, Grafana URL, and kubeconfig command
🔍 STACKIT Cloud Advisor

1. Identified STACKIT Services

Based on the Terraform configuration, the following services are being provisioned or utilized:

  • SKE (STACKIT Kubernetes Engine): A managed Kubernetes cluster is being created (stackit_ske_cluster), including a node pool using the flatcar OS and enabling Kubernetes API audit logging (audit.enabled = true).
  • Network: A dedicated STACKIT network (stackit_network) is being provisioned to host the SKE nodes.
  • Observability: A managed observability instance (stackit_observability_instance) is being provisioned to act as the long-term storage and visualization layer (Loki/Grafana). It also includes the creation of credentials (stackit_observability_credential) for ingestion.
  • Telemetry Router: A managed instance (stackit_telemetryrouter_instance) is being provisioned to act as the central ingestion point. It includes an access token (stackit_telemetryrouter_access_token) and a destination (stackit_telemetryrouter_destination) configured to forward logs to the Observability service via the OTLP protocol.
  • Telemetry Link: A link (stackit_telemetrylink) is being established to connect the STACKIT project to the Telemetry Router, enabling the automatic capture of audit streams.

Proposed Data Flow Architecture:

[ SKE Cluster ] --(Audit Logs)--> [ Telemetry Link ]
                                        |
                                        v
[ Observability ] <--(OTLP)--- [ Telemetry Router ]
 (Loki/Grafana)

2. STACKIT Best Practices & Architectural Review

While the PR follows the functional requirements for a demo, I have identified several points regarding production readiness and best practices:

  • Hub-and-Spoke Topology: The code currently deploys the Telemetry Router in the same project as the SKE cluster. As noted in the code comments, for production environments, it is a best practice to run a single Telemetry Router in a dedicated hub project to aggregate telemetry from multiple "spoke" projects Architecture.
  • Telemetry Link Necessity: The implementation correctly uses a stackit_telemetrylink. This is a requirement to ensure that STACKIT internal logs are actually redirected to your configured destinations Architecture.
  • Resource Lifecycle: The use of lifecycle { ignore_changes = [...] } on the SKE cluster is a good practice here, as it prevents Terraform from attempting to revert automatic Kubernetes version upgrades or node pool scaling performed by the STACKIT control plane.
  • Security (ACLs): The telemetry_acl is currently set to ["0.0.0.0/0"]. For a production observability instance, this should be restricted to specific known IP ranges to minimize the attack surface.

3. Quotas, Regional Constraints, and Limitations

Please be aware of the following constraints when deploying this configuration:

  • Regionality Constraint (Critical): The Telemetry Link assumes that the configured Telemetry Router is located in the same region as the resource being linked Architecture. Ensure var.stackit_region is consistent across all components.
  • Telemetry Router Scope: The Telemetry Router is currently designed exclusively to handle Audit Logs. It cannot be used to ingest custom application logs, metrics, or traces Overview.
  • Router Placement: Due to technical limitations, a Telemetry Router must always be located within a project Architecture.
  • Availability: The Telemetry Router service is currently Generally Available (GA) in the EU01 (Germany) and EU02 (Austria) regions Overview.
  • Destination URI Limit: If you modify the stackit_telemetryrouter_destination configuration, note that the uri parameter has a limit of 1024 characters Create And Manage Destinations.
  • Observability Plan: The configuration uses Observability-Large-EU01. Ensure your project has the necessary quota/permissions for this specific plan size.
⚠️ 🔒 Security Review
🤖 STACKIT Model Serving

✅ No security issues found.

🔍 STACKIT Cloud Advisor

1. STACKIT IAM or Authorization Misconfigurations

  • 🟡 Medium: Potential Over-Privileged Telemetry Link
    The stackit_telemetrylink resource is configured with resource_type = "project" and points to the entire var.stackit_project_id [060-telemetry-link.tf]. While this is necessary to ingest logs from the project, ensure that the service account executing this Terraform plan follows the principle of least privilege, as it requires permissions to link entire projects to the Telemetry Router.

2. Missing STACKIT-Specific Security Controls

  • 🔴 High: Overly Permissive Observability ACL
    In 020-variables.tf, the telemetry_acl variable defaults to ["0.0.0.0/0"]. This is applied to the stackit_observability_instance in 040-observability.tf.

    • Risk: By default, Observability is available on the Internet and accessible to everyone How To Control Instance Access. Using 0.0.0.0/0 allows any IP address to attempt access to your Grafana, Metrics, Logs, and Traces endpoints How To Control Instance Access.
    • Recommendation: Restrict the ACL to known corporate IP ranges or specific administrative IPs. For production, avoid 0.0.0.0/0 Enhance The Security Of Your Cluster. You can use the stackit_public_ip_ranges data source to at least allow all STACKIT internal services while restricting external access Public Ip Ranges.
  • 🔴 High: Missing SKE API Server Access Control (ACL)
    The stackit_ske_cluster resource in 070-ske-cluster.tf does not define the extensions.acl block.

  • 🟡 Medium: Lack of Network Policies for SKE
    The diff defines a stackit_network for the nodes, but there are no Kubernetes NetworkPolicy resources defined to control pod-to-pod traffic [030-network.tf].

    • Recommendation: Implement a "Default Deny" policy for all namespaces and then explicitly allow necessary traffic to minimize the blast radius of a potential workload compromise Enhance The Security Of Your Cluster.

3. Secrets and Sensitive Values

  • 🟢 Low: Service Account Key Path
    The provider configuration uses service_account_key_path = var.stackit_service_account_key_path [010-provider.tf]. While the path itself is not a secret, the security of the deployment relies entirely on the protection of the JSON file located at that path. Ensure this file is never committed to version control.

4. STACKIT Compliance and Audit-Logging Considerations

  • 🟢 Low: Audit Logging Implementation
    The configuration correctly enables audit = { enabled = true } for the SKE cluster [070-ske-cluster.tf] and sets up a pipeline via Telemetry Router to Observability. This aligns with best practices for centralized logging and auditability Enhance The Security Of Your Cluster.

Summary of Security Posture

Component Finding Severity Mitigation
Observability Publicly accessible via 0.0.0.0/0 🔴 High Restrict acl to specific CIDRs How To Control Instance Access.
SKE Cluster API Server lacks network ACL 🔴 High Enable extensions.acl in the cluster resource Configure Access Control For Your Cluster.
SKE Cluster Missing Pod Network Policies 🟡 Medium Implement NetworkPolicy for pod isolation Enhance The Security Of Your Cluster.
IAM Project-wide Telemetry Link 🟡 Medium Audit service account permissions for the link operation.

Architectural Topology of the Proposed Change:

[ Internet ] 
     |
     | (Blocked by ACL if configured)
     v
[ STACKIT Observability ] <--- (Ingest via OTLP) --- [ Telemetry Router ]
                                                           ^
                                                           | (Telemetry Link)
                                                           |
[ SKE Cluster ] -------------------------------------------+
     |
     +-- [ API Server ] (Needs ACL Extension)
     |
     +-- [ Worker Nodes ] (In stackit_network)
✅ 📐 Example Consistency
🤖 STACKIT Model Serving

✅ Example follows repository conventions.

🔍 STACKIT Cloud Advisor

The review of the provided git diff against the repository conventions and STACKIT provider standards is complete. While the example is structurally sound and follows most naming and documentation patterns, there are specific deviations regarding provider versioning and STACKIT-specific configuration requirements.

🚩 Deviations from Repository Conventions

1. Missing Variable Descriptions

The convention requires all variables to have a description attribute. In 020-variables.tf, the stackit_project_id variable is missing this attribute.

Fix:

variable "stackit_project_id" {
  type        = string
  default     = "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
  description = "The STACKIT Project ID where resources will be provisioned."
}

2. Non-Explicit Provider Version Constraints

The convention states that all providers in required_providers blocks must have an explicit version constraint. In 010-provider.tf, the versions use the >= operator, which allows for minor/patch updates but is technically a range constraint rather than a fixed or strictly pinned version. To strictly adhere to "explicit version constraint" in a production-grade repository, it is safer to pin to a specific version or a more restrictive range.

Fix:

terraform {
  required_providers {
    stackit = {
      source  = "stackitcloud/stackit"
      version = "0.113.0" # Pinning to the specific version used in the lock file
    }
    kubernetes = {
      source  = "hashicorp/kubernetes"
      version = "3.2.1"
    }
  }
}

🛠️ STACKIT-Specific Advisory

Beta Resource Configuration

The example correctly implements the requirement for using beta resources. In 010-provider.tf, the enable_beta_resources = true flag is set within the stackit provider block. This is mandatory when interacting with experimental or preview features, such as the SKE audit logging feature mentioned in the README.md.

Architecture Overview

The example implements a sophisticated telemetry pipeline. Below is the logical flow of the data as defined in the Terraform resources:

[ SKE Cluster ] 
      |
      | (Audit Logs via Private Preview)
      v
[ Telemetry Router ] <--- [ Telemetry Link ] <--- [ Project ID ]
      |
      | (Filtered via Destination Config)
      v
[ Observability Instance ] (Loki + Grafana)

Provider Usage Note

The use of ephemeral "stackit_ske_kubeconfig" "this" is a highly efficient pattern for managing the lifecycle of Kubernetes credentials within a single Terraform run, ensuring that the kubernetes provider can immediately interact with the newly created cluster without manual intervention.

Summary Table

Requirement Status Note
3-digit numeric prefixes ✅ Pass Files use 010-, 020-, etc.
README.md & MAINTAINERS.md ✅ Pass Both files are present.
Variable description & snake_case ❌ Fail stackit_project_id is missing a description.
Explicit provider versions ⚠️ Warning Uses >= instead of pinned versions.
.terraform.lock.hcl committed ✅ Pass File is present in the diff.
Apache 2.0 headers ✅ Pass Present on all .tf files.
⚠️ 📚 Example README
🤖 STACKIT Model Serving

✅ Example READMEs are complete.

🔍 STACKIT Cloud Advisor

Review of examples/ske-kubeapi-audit-log

I have reviewed the provided git diff for the new example directory. Below is my architectural assessment regarding naming conventions and documentation quality.

1. Naming Convention Analysis

The directory name ske-kubeapi-audit-log is compliant with our standards.

  • Clarity: It explicitly identifies the primary service involved (SKE - STACKIT Kubernetes Engine) and the specific technical use-case (kubeapi-audit-log).
  • Accuracy: The content of the Terraform files (specifically 070-ske-cluster.tf) confirms that the core objective is enabling the audit = { enabled = true } feature on an SKE cluster.
  • Consistency: The name follows the established pattern of [service]-[use-case] seen in other examples (e.g., ske-velero-backup).

2. README Quality Assessment

The README.md is high quality and meets all required criteria for a production-ready example.

  • Service Identification: It clearly outlines the orchestration of multiple STACKIT services, including SKE, Telemetry Router, and Observability (Loki + Grafana).
  • Demonstration Scope: It provides a deep dive into the "Layered Filtering" logic (Source \rightarrow Router Instance \rightarrow Destination \rightarrow Query time), which is critical for users to understand how to manage the high volume of audit logs.
  • Usage Section: A clear "How to use" section is provided with the necessary terraform init and terraform apply commands.
  • Technical Depth:
    • It includes a detailed table of resource mappings.
    • It provides specific LogQL queries for practical testing.
    • It includes a crucial warning regarding the Private Preview status of the SKE audit feature, advising users to contact STACKIT for enablement.

Architectural Topology Overview

The example demonstrates the following data flow:

[ SKE Cluster ] 
      |
      | (Audit Logs via Telemetry Link)
      v
[ Telemetry Router ] 
      |
      | (Filtered via Destination Config)
      v
[ Observability Instance ]
      |
      | (Visualized via Grafana)
      v
[ User / Admin ]

Verdict:
The example is well-structured, accurately named, and provides sufficient documentation for a user to successfully deploy and validate the audit logging pipeline.

✅ 💬 Commit Messages
🤖 STACKIT Model Serving

✅ Commit messages are descriptive.

🔍 STACKIT Cloud Advisor

All provided commit messages are sufficiently descriptive.

✅ Commit messages are descriptive.


Generated automatically — treat as a hint, not a gate.

## 🤖 AI PR Review > [`75c59515`](https://professional-service.git.onstackit.cloud/professional-service-best-practices/professional-service/commit/75c5951542eb34e28fe72585dbccee82eae1eab6) · STACKIT Model Serving & STACKIT Cloud Advisor <details> <summary>⚠️ 📝 Spelling & Grammar</summary> <details> <summary>🤖 STACKIT Model Serving</summary> - `examples/ske-kubeapi-audit-log/README.md:15` — “comming” → “coming” - `examples/ske-kubeapi-audit-log/README.md:15` — “overwhleming” → “overwhelming” - `examples/ske-kubeapi-audit-log/README.md:24` — “review” → “preview” (in “private review” → should be “private preview” to match context) - `examples/ske-kubeapi-audit-log/README.md:59` — “STACKIT Logs only focusing on have a managed Loki” → “STACKIT Logs focuses only on providing a managed Loki” - `examples/ske-kubeapi-audit-log/README.md:60` — “Observability comes with a managed Grafana on top.” → “Observability includes a managed Grafana on top.” (for smoother phrasing) - `examples/ske-kubeapi-audit-log/README.md:81` — “The resource attributes STACKIT sets on every forwarded record (dots become underscores once they arrive as Loki labels):” → “The resource attributes STACKIT sets on every forwarded record (dots become underscores when they arrive as Loki labels):” (more natural phrasing) </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### Spelling and Grammar Review I have reviewed the provided git diff for spelling and grammar errors in the prose content, including Markdown documentation, Terraform variable descriptions, and inline comments. I have also verified the usage of STACKIT product names and technical terminology against the provided documentation. The following errors were identified: | File | Line | Wrong Text | Correction | | :--- | :--- | :--- | :--- | | `examples/ske-kubeapi-audit-log/010-provider.tf` | 18 | `distributed under an "AS IS" BASIS` | `distributed on an "AS IS" BASIS` | | `examples/ske-kubeapi-audit-log/020-variables.tf` | 64 | `Access control list of the Observability/Logs.` | `Access control list of the Observability/Logs.` (Note: While grammatically acceptable, "for" is more idiomatic: `Access control list for Observability/Logs.`) | | `examples/ske-kubeapi-audit-log/080-canary-workload.tf` | 13 | `into the audit` | `into the audit stream` (Sentence is truncated/incomplete) | | `examples/ske-kubeapi-audit-log/README.md` | 11 | `comming from kublet or gardener` | `coming from kubelet or gardener` | | `examples/ske-kubeapi-audit-log/README.md` | 11 | `This can be overwhleming` | `This can be overwhelming` | | `examples/ske-kubeapi-audit-log/README.md` | 15 | `in private review` | `in private preview` | | `examples/ske-kubeapi-audit-log/README.md` | 53 | `STACKIT Logs only focusing on have a managed Loki` | `STACKIT Logs only focus on having a managed Loki` | | `examples/ske-kubeapi-audit-log/README.md` | 108 | `where the kublet/gardener noise` | `where the kubelet/gardener noise` | </details> </details> <details> <summary>⚠️ 🏗️ Infrastructure Changes</summary> <details> <summary>🤖 STACKIT Model Serving</summary> - Creates STACKIT network for SKE nodes with IPv4 prefix and nameservers - Creates STACKIT Observability instance (Loki+Grafana) with log retention and ACL - Creates STACKIT Telemetry Router instance with access token and destination config - Creates STACKIT Telemetry Link to route project audit logs to the router - Creates SKE Kubernetes cluster with audit enabled, using Flatcar OS and node pools - Creates ephemeral kubeconfig for cluster access during Terraform apply - Deploys Kubernetes namespace and ConfigMap as audit canary workload via ephemeral kubeconfig - Exposes outputs for cluster name, router IDs, Grafana URL, and kubeconfig command </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### 1. Identified STACKIT Services Based on the Terraform configuration, the following services are being provisioned or utilized: * **SKE (STACKIT Kubernetes Engine):** A managed Kubernetes cluster is being created (`stackit_ske_cluster`), including a node pool using the `flatcar` OS and enabling Kubernetes API audit logging (`audit.enabled = true`). * **Network:** A dedicated STACKIT network (`stackit_network`) is being provisioned to host the SKE nodes. * **Observability:** A managed observability instance (`stackit_observability_instance`) is being provisioned to act as the long-term storage and visualization layer (Loki/Grafana). It also includes the creation of credentials (`stackit_observability_credential`) for ingestion. * **Telemetry Router:** A managed instance (`stackit_telemetryrouter_instance`) is being provisioned to act as the central ingestion point. It includes an access token (`stackit_telemetryrouter_access_token`) and a destination (`stackit_telemetryrouter_destination`) configured to forward logs to the Observability service via the OTLP protocol. * **Telemetry Link:** A link (`stackit_telemetrylink`) is being established to connect the STACKIT project to the Telemetry Router, enabling the automatic capture of audit streams. **Proposed Data Flow Architecture:** ```text [ SKE Cluster ] --(Audit Logs)--> [ Telemetry Link ] | v [ Observability ] <--(OTLP)--- [ Telemetry Router ] (Loki/Grafana) ``` ### 2. STACKIT Best Practices & Architectural Review While the PR follows the functional requirements for a demo, I have identified several points regarding production readiness and best practices: * **Hub-and-Spoke Topology:** The code currently deploys the Telemetry Router in the same project as the SKE cluster. As noted in the code comments, for production environments, it is a best practice to run a single Telemetry Router in a **dedicated hub project** to aggregate telemetry from multiple "spoke" projects [Architecture](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/basics/architecture/). * **Telemetry Link Necessity:** The implementation correctly uses a `stackit_telemetrylink`. This is a requirement to ensure that STACKIT internal logs are actually redirected to your configured destinations [Architecture](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/basics/architecture/). * **Resource Lifecycle:** The use of `lifecycle { ignore_changes = [...] }` on the SKE cluster is a good practice here, as it prevents Terraform from attempting to revert automatic Kubernetes version upgrades or node pool scaling performed by the STACKIT control plane. * **Security (ACLs):** The `telemetry_acl` is currently set to `["0.0.0.0/0"]`. For a production observability instance, this should be restricted to specific known IP ranges to minimize the attack surface. ### 3. Quotas, Regional Constraints, and Limitations Please be aware of the following constraints when deploying this configuration: * **Regionality Constraint (Critical):** The **Telemetry Link** assumes that the configured **Telemetry Router is located in the same region** as the resource being linked [Architecture](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/basics/architecture/). Ensure `var.stackit_region` is consistent across all components. * **Telemetry Router Scope:** The Telemetry Router is currently designed exclusively to handle **Audit Logs**. It cannot be used to ingest custom application logs, metrics, or traces [Overview](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/basics/overview/). * **Router Placement:** Due to technical limitations, a Telemetry Router must always be located within a project [Architecture](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/basics/architecture/). * **Availability:** The Telemetry Router service is currently Generally Available (GA) in the **EU01** (Germany) and **EU02** (Austria) regions [Overview](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/basics/overview/). * **Destination URI Limit:** If you modify the `stackit_telemetryrouter_destination` configuration, note that the `uri` parameter has a limit of **1024 characters** [Create And Manage Destinations](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/how-tos/create-and-manage-destinations/). * **Observability Plan:** The configuration uses `Observability-Large-EU01`. Ensure your project has the necessary quota/permissions for this specific plan size. </details> </details> <details> <summary>⚠️ 🔒 Security Review</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ No security issues found. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### 1. STACKIT IAM or Authorization Misconfigurations * 🟡 **Medium: Potential Over-Privileged Telemetry Link** The `stackit_telemetrylink` resource is configured with `resource_type = "project"` and points to the entire `var.stackit_project_id` [060-telemetry-link.tf]. While this is necessary to ingest logs from the project, ensure that the service account executing this Terraform plan follows the principle of least privilege, as it requires permissions to link entire projects to the **Telemetry Router**. ### 2. Missing STACKIT-Specific Security Controls * 🔴 **High: Overly Permissive Observability ACL** In `020-variables.tf`, the `telemetry_acl` variable defaults to `["0.0.0.0/0"]`. This is applied to the `stackit_observability_instance` in `040-observability.tf`. * **Risk:** By default, **Observability** is available on the Internet and accessible to everyone [How To Control Instance Access](https://docs.stackit.cloud/products/logging-and-monitoring/observability/how-tos/how-to-control-instance-access/). Using `0.0.0.0/0` allows any IP address to attempt access to your **Grafana**, **Metrics**, **Logs**, and **Traces** endpoints [How To Control Instance Access](https://docs.stackit.cloud/products/logging-and-monitoring/observability/how-tos/how-to-control-instance-access/). * **Recommendation:** Restrict the ACL to known corporate IP ranges or specific administrative IPs. For production, avoid `0.0.0.0/0` [Enhance The Security Of Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/enhance-the-security-of-your-cluster/). You can use the `stackit_public_ip_ranges` data source to at least allow all STACKIT internal services while restricting external access [Public Ip Ranges](https://registry.terraform.io/providers/stackitcloud/stackit/latest/docs/data-sources/public_ip_ranges). * 🔴 **High: Missing SKE API Server Access Control (ACL)** The `stackit_ske_cluster` resource in `070-ske-cluster.tf` does not define the `extensions.acl` block. * **Risk:** Without the `acl` extension enabled and configured with `allowedCidrs`, the Kubernetes API server is not protected by the additional layer of STACKIT network isolation [Configure Access Control For Your Cluster](https://docs.stackit.cloud/de/products/runtime/kubernetes-engine/how-tos/configure-access-control-for-your-cluster/) [Configure Access Control For Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/configure-access-control-for-your-cluster/). * **Recommendation:** Implement the `acl` extension within the SKE cluster configuration to restrict API access to specific, trusted CIDR ranges [Configure Access Control For Your Cluster](https://docs.stackit.cloud/de/products/runtime/kubernetes-engine/how-tos/configure-access-control-for-your-cluster/). * 🟡 **Medium: Lack of Network Policies for SKE** The diff defines a `stackit_network` for the nodes, but there are no Kubernetes `NetworkPolicy` resources defined to control pod-to-pod traffic [030-network.tf]. * **Recommendation:** Implement a "Default Deny" policy for all namespaces and then explicitly allow necessary traffic to minimize the blast radius of a potential workload compromise [Enhance The Security Of Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/enhance-the-security-of-your-cluster/). ### 3. Secrets and Sensitive Values * 🟢 **Low: Service Account Key Path** The provider configuration uses `service_account_key_path = var.stackit_service_account_key_path` [010-provider.tf]. While the path itself is not a secret, the security of the deployment relies entirely on the protection of the JSON file located at that path. Ensure this file is never committed to version control. ### 4. STACKIT Compliance and Audit-Logging Considerations * 🟢 **Low: Audit Logging Implementation** The configuration correctly enables `audit = { enabled = true }` for the **SKE** cluster [070-ske-cluster.tf] and sets up a pipeline via **Telemetry Router** to **Observability**. This aligns with best practices for centralized logging and auditability [Enhance The Security Of Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/enhance-the-security-of-your-cluster/). --- ### Summary of Security Posture | Component | Finding | Severity | Mitigation | | :--- | :--- | :--- | :--- | | **Observability** | Publicly accessible via `0.0.0.0/0` | 🔴 High | Restrict `acl` to specific CIDRs [How To Control Instance Access](https://docs.stackit.cloud/products/logging-and-monitoring/observability/how-tos/how-to-control-instance-access/). | | **SKE Cluster** | API Server lacks network ACL | 🔴 High | Enable `extensions.acl` in the cluster resource [Configure Access Control For Your Cluster](https://docs.stackit.cloud/de/products/runtime/kubernetes-engine/how-tos/configure-access-control-for-your-cluster/). | | **SKE Cluster** | Missing Pod Network Policies | 🟡 Medium | Implement `NetworkPolicy` for pod isolation [Enhance The Security Of Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/enhance-the-security-of-your-cluster/). | | **IAM** | Project-wide Telemetry Link | 🟡 Medium | Audit service account permissions for the link operation. | **Architectural Topology of the Proposed Change:** ```text [ Internet ] | | (Blocked by ACL if configured) v [ STACKIT Observability ] <--- (Ingest via OTLP) --- [ Telemetry Router ] ^ | (Telemetry Link) | [ SKE Cluster ] -------------------------------------------+ | +-- [ API Server ] (Needs ACL Extension) | +-- [ Worker Nodes ] (In stackit_network) ``` </details> </details> <details> <summary>✅ 📐 Example Consistency</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ Example follows repository conventions. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> The review of the provided git diff against the repository conventions and STACKIT provider standards is complete. While the example is structurally sound and follows most naming and documentation patterns, there are specific deviations regarding provider versioning and STACKIT-specific configuration requirements. ### 🚩 Deviations from Repository Conventions #### 1. Missing Variable Descriptions The convention requires **all** variables to have a `description` attribute. In `020-variables.tf`, the `stackit_project_id` variable is missing this attribute. **Fix:** ```hcl variable "stackit_project_id" { type = string default = "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx" description = "The STACKIT Project ID where resources will be provisioned." } ``` #### 2. Non-Explicit Provider Version Constraints The convention states that all providers in `required_providers` blocks must have an **explicit version constraint**. In `010-provider.tf`, the versions use the `>=` operator, which allows for minor/patch updates but is technically a range constraint rather than a fixed or strictly pinned version. To strictly adhere to "explicit version constraint" in a production-grade repository, it is safer to pin to a specific version or a more restrictive range. **Fix:** ```hcl terraform { required_providers { stackit = { source = "stackitcloud/stackit" version = "0.113.0" # Pinning to the specific version used in the lock file } kubernetes = { source = "hashicorp/kubernetes" version = "3.2.1" } } } ``` ### 🛠️ STACKIT-Specific Advisory #### Beta Resource Configuration The example correctly implements the requirement for using beta resources. In `010-provider.tf`, the `enable_beta_resources = true` flag is set within the `stackit` provider block. This is mandatory when interacting with experimental or preview features, such as the SKE audit logging feature mentioned in the `README.md`. #### Architecture Overview The example implements a sophisticated telemetry pipeline. Below is the logical flow of the data as defined in the Terraform resources: ```text [ SKE Cluster ] | | (Audit Logs via Private Preview) v [ Telemetry Router ] <--- [ Telemetry Link ] <--- [ Project ID ] | | (Filtered via Destination Config) v [ Observability Instance ] (Loki + Grafana) ``` #### Provider Usage Note The use of `ephemeral "stackit_ske_kubeconfig" "this"` is a highly efficient pattern for managing the lifecycle of Kubernetes credentials within a single Terraform run, ensuring that the `kubernetes` provider can immediately interact with the newly created cluster without manual intervention. ### Summary Table | Requirement | Status | Note | | :--- | :--- | :--- | | 3-digit numeric prefixes | ✅ Pass | Files use `010-`, `020-`, etc. | | README.md & MAINTAINERS.md | ✅ Pass | Both files are present. | | Variable `description` & `snake_case` | ❌ Fail | `stackit_project_id` is missing a description. | | Explicit provider versions | ⚠️ Warning | Uses `>=` instead of pinned versions. | | `.terraform.lock.hcl` committed | ✅ Pass | File is present in the diff. | | Apache 2.0 headers | ✅ Pass | Present on all `.tf` files. | </details> </details> <details> <summary>⚠️ 📚 Example README</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ Example READMEs are complete. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### Review of `examples/ske-kubeapi-audit-log` I have reviewed the provided git diff for the new example directory. Below is my architectural assessment regarding naming conventions and documentation quality. #### 1. Naming Convention Analysis The directory name `ske-kubeapi-audit-log` is **compliant** with our standards. * **Clarity:** It explicitly identifies the primary service involved (**SKE** - STACKIT Kubernetes Engine) and the specific technical use-case (**kubeapi-audit-log**). * **Accuracy:** The content of the Terraform files (specifically `070-ske-cluster.tf`) confirms that the core objective is enabling the `audit = { enabled = true }` feature on an SKE cluster. * **Consistency:** The name follows the established pattern of `[service]-[use-case]` seen in other examples (e.g., `ske-velero-backup`). #### 2. README Quality Assessment The `README.md` is **high quality** and meets all required criteria for a production-ready example. * **Service Identification:** It clearly outlines the orchestration of multiple STACKIT services, including **SKE**, **Telemetry Router**, and **Observability** (Loki + Grafana). * **Demonstration Scope:** It provides a deep dive into the "Layered Filtering" logic (Source $\rightarrow$ Router Instance $\rightarrow$ Destination $\rightarrow$ Query time), which is critical for users to understand how to manage the high volume of audit logs. * **Usage Section:** A clear "How to use" section is provided with the necessary `terraform init` and `terraform apply` commands. * **Technical Depth:** * It includes a detailed table of resource mappings. * It provides specific **LogQL** queries for practical testing. * It includes a crucial warning regarding the **Private Preview** status of the SKE audit feature, advising users to contact STACKIT for enablement. #### Architectural Topology Overview The example demonstrates the following data flow: ```text [ SKE Cluster ] | | (Audit Logs via Telemetry Link) v [ Telemetry Router ] | | (Filtered via Destination Config) v [ Observability Instance ] | | (Visualized via Grafana) v [ User / Admin ] ``` **Verdict:** The example is well-structured, accurately named, and provides sufficient documentation for a user to successfully deploy and validate the audit logging pipeline. </details> </details> <details> <summary>✅ 💬 Commit Messages</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ Commit messages are descriptive. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> All provided commit messages are sufficiently descriptive. ✅ Commit messages are descriptive. </details> </details> --- _Generated automatically — treat as a hint, not a gate._

Be aware: pre-commit failed and the AI Agent found some grammar issues

Be aware: pre-commit failed and the AI Agent found some grammar issues
mauritz.uphoff force-pushed example/ske-kubeapi-audit-log from 75c5951542
Some checks failed
Default CI / Check for Open TODOs (pull_request) Successful in 1m40s
AI PR Review / AI PR Review (pull_request) Successful in 3m29s
Default CI / Pre-Commit Hooks (pull_request) Failing after 11m35s
to c103c36151
Some checks failed
Default CI / Check for Open TODOs (pull_request) Successful in 43s
Default CI / Check README Tags (pull_request) Failing after 1m5s
Default CI / Check AGENTS.md is up to date (pull_request) Has been skipped
AI PR Review / AI PR Review (pull_request) Successful in 2m7s
Default CI / Pre-Commit Hooks (pull_request) Failing after 2m14s
2026-08-26 15:11:26 +00:00
Compare

🤖 AI PR Review

c103c361 · STACKIT Model Serving & STACKIT Cloud Advisor

⚠️ 📝 Spelling & Grammar
🤖 STACKIT Model Serving
  • examples/ske-kubeapi-audit-log/README.md:12 — “comming” → “coming”
  • examples/ske-kubeapi-audit-log/README.md:13 — “overwhleming” → “overwhelming”
  • examples/ske-kubeapi-audit-log/README.md:20 — “review” → “preview”
  • examples/ske-kubeapi-audit-log/README.md:75 — “have a managed Loki” → “has a managed Loki”
🔍 STACKIT Cloud Advisor

Spelling and Grammar Review

I have reviewed the provided git diff for spelling and grammar errors in the prose content, including Markdown documentation, Terraform variable descriptions, and inline comments. I have also verified the usage of STACKIT product names and technical terminology against the provided documentation.

The following errors were identified:

File: examples/ske-kubeapi-audit-log/010-provider.tf

  • Line 18: distributed under the License is distributed on an "AS IS" BASIS, \rightarrow distributed under the License is distributed on an "AS IS" basis,
    (Note: While this is part of a standard Apache License header, "BASIS" is capitalized in a way that deviates from standard prose, though often accepted in legal headers. If strictly following grammar rules for prose, it should be lowercase.)

File: examples/ske-kubeapi-audit-log/README.md

  • Line 10: doing get/list/watch calls comming from kublet or gardener \rightarrow doing get/list/watch calls coming from kubelet or gardener
  • Line 10: This can be overwhleming so this example aims to show how you can view relevant audit logs as well. \rightarrow This can be overwhelming, so this example aims to show how you can view relevant audit logs as well.
  • Line 16: because this feature is in private review \rightarrow because this feature is in private preview
  • Line 54: Both are Loki underneath. STACKIT Logs only focusing on have a managed Loki and Observability comes with a managed Grafana on top. \rightarrow Both are Loki underneath. STACKIT Logs only focuses on having a managed Loki, and Observability comes with a managed Grafana on top.
  • Line 138: The apply already produces findable audit records: the canary Namespace and ConfigMap are created \rightarrow The apply already produces findable audit records: the canary namespace and ConfigMap are created.

Architect's Note on Terminology:
The usage of SKE (STACKIT Kubernetes Engine), Telemetry Router, Observability, and Telemetry Link is consistent with STACKIT product naming conventions. The technical distinction between STACKIT Logs and STACKIT Observability is correctly maintained in the prose.

⚠️ 🏗️ Infrastructure Changes
🤖 STACKIT Model Serving
  • Creates STACKIT network for SKE nodes with IPv4 prefix and nameservers
  • Creates STACKIT Observability instance (Loki + Grafana) with log retention and ACL
  • Creates STACKIT Telemetry Router instance and associated access token
  • Creates STACKIT Telemetry Link to route project logs to the router
  • Creates STACKIT SKE Kubernetes cluster with audit enabled, node pools, and network attachment
  • Creates ephemeral kubeconfig for cluster access during Terraform apply
  • Deploys Kubernetes namespace and config map via ephemeral kubeconfig as audit canary
  • Exposes outputs for cluster name, router ID/URI, link ID, Grafana URL, and kubeconfig command

⚠️ No destructive changes detected — all resources are newly created. No force-replacements or deletions in this diff.

🔍 STACKIT Cloud Advisor

1. Identified STACKIT Services

Based on the Terraform configuration, the following services are being provisioned or utilized:

  • SKE (STACKIT Kubernetes Engine): A cluster is being provisioned (stackit_ske_cluster) with audit logging enabled. It uses a custom network and a specific node pool configuration.
  • Network: A dedicated STACKIT network (stackit_network) is being created to host the SKE nodes.
  • Observability: A managed instance (stackit_observability_instance) is being provisioned using the Observability-Large-EU01 plan to store and visualize logs.
  • Telemetry Router: A managed router instance (stackit_telemetryrouter_instance) is being deployed to act as the ingestion point for audit logs.
  • Telemetry Link: A link (stackit_telemetrylink) is being established to connect the STACKIT project to the Telemetry Router.

Data Flow Architecture:

[ SKE Cluster ] --(Audit Logs)--> [ Telemetry Link ]
                                         |
                                         v
[ Observability ] <--(OTLP/Auth)-- [ Telemetry Router ]

2. STACKIT Best Practices & Architectural Review

Security & Access Control

  • Telemetry ACL: The variable telemetry_acl is defaulted to ["0.0.0.0/0"]. Recommendation: For production environments, restrict this to specific known IP ranges to minimize the attack surface of your Observability instance.
  • Router Placement: The code includes a comment noting that for demo purposes, the router lives in the same project as the cluster. Best Practice: In a production-grade architecture, you should deploy the Telemetry Router in a dedicated "Hub" project to centralize audit log collection from multiple "Spoke" projects Architecture.
  • Credential Management: The use of stackit_observability_credential to pass technical credentials to the router is the correct way to handle OTLP authentication via Basic Auth.

Observability Configuration

  • Log Retention: The observability_logs_retention_days is set to 7.
    • Note: The Observability-Large plan supports a default of 5 days and can be configured up to 30 days via the API Service Plans Observability. Ensure your retention requirement aligns with your compliance needs.
  • Filtering: The stackit_telemetryrouter_destination uses a specific filter for stackit.log.kind = "kubernetes-audit". This is an excellent practice to prevent "noise" and unnecessary costs by only routing relevant audit data to the Observability instance.

3. Quotas, Regional Constraints, and Limitations

Regional Constraints

  • Telemetry Router Availability: The Telemetry Router is currently only available in the EU01 (Germany) and EU02 (Austria) regions Scope and Limitations.
  • Telemetry Link Locality: A critical constraint is that the Telemetry Link requires the Telemetry Router to be in the same region as the resources being linked Architecture. Since your stackit_region defaults to eu01, ensure your cluster and router remain in this region.

Service Limitations

  • Telemetry Router Data Type: The router is currently specialized for Audit Logs only. You cannot use this specific router instance to send custom application metrics or traces Scope and Limitations.
  • Observability Plan Limits: You are using the Observability-Large plan. Be aware of the following hard limits for this plan:
  • Observability Backup: Note that while configuration and dashboards are backed up automatically, there is no backup of the telemetry data itself (metrics, logs, or traces) Service Certificate.
⚠️ 🔒 Security Review
🤖 STACKIT Model Serving

✅ No security issues found.

🔍 STACKIT Cloud Advisor

1. STACKIT IAM or Authorization Misconfigurations

  • 🟡 Medium: Telemetry Router Access Token Scope
    The stackit_telemetryrouter_access_token is created to allow the stackit_telemetrylink to subscribe the project to the router. While this is the functional requirement for the link, ensure that the principle of least privilege is applied to the identity using this token if it were to be exported or used outside of the automated Terraform flow.

2. Missing STACKIT-Specific Security Controls

  • 🔴 High: Overly Permissive Observability ACL
    In 020-variables.tf, the telemetry_acl variable defaults to ["0.0.0.0/0"]. The Observability service is available on the Internet by default How To Control Instance Access. Using 0.0.0.0/0 allows full access to Grafana, Metrics, Logs, and Traces from any IP address How To Control Instance Access.

    • Recommendation: Restrict this to specific corporate IP ranges or use the stackit_public_ip_ranges data source to at least limit access to known STACKIT service ranges Public Ip Ranges.
  • 🔴 High: Missing SKE API Server Access Control (ACL)
    The stackit_ske_cluster resource in 070-ske-cluster.tf does not define the extensions.acl block. Without this, the Kubernetes API server is not protected by the SKE ACL functionality, which is a critical layer of security for limiting access to the API Enhance The Security Of Your Cluster Configure Access Control For Your Cluster.

  • 🟡 Medium: Lack of Network Policies
    The current configuration defines a stackit_network for the nodes, but there are no Kubernetes NetworkPolicy resources defined to control pod-to-pod traffic Enhance The Security Of Your Cluster.

    • Recommendation: Implement a "Default Deny" policy for the audit-canary namespace and subsequent production namespaces to enforce granular traffic control Enhance The Security Of Your Cluster.
  • 🟡 Medium: Absence of Pod Security Standards (PSS)
    The audit-canary namespace is created without any Pod Security Standard labels Enhance The Security Of Your Cluster.

    • Recommendation: Apply pod-security.kubernetes.io/enforce: restricted labels to namespaces to prevent privileged containers and enforce non-root users Enhance The Security Of Your Cluster.

3. Secrets and Sensitive Values

  • 🟢 Low: Service Account Key Path
    The variable stackit_service_account_key_path is used to point to a JSON key file. While the path itself is not a secret, the security of the entire deployment relies on the protection of the file at that path. Ensure this file is never committed to version control.

4. STACKIT Compliance and Audit-Logging Considerations

  • 🟢 Low: Audit Log Retention
    The observability_logs_retention_days is set to 7 days [020-variables.tf]. While this is sufficient for a "demo" (as suggested by the code comments), production compliance requirements often mandate longer retention periods for audit logs.
  • 🟢 Low: Telemetry Router Architecture
    The code notes that the router lives in the same project as the cluster for demo purposes [050-telemetry-router.tf]. For production environments, a more secure architecture involves running the Telemetry Router in a dedicated "hub" project to centralize logging and isolate the observability infrastructure from workload projects.
[ Project: Workload ]          [ Project: Hub/Logging ]
+-------------------+          +-----------------------+
| SKE Cluster       |          | Telemetry Router      |
| (Audit Logs) ---->|--------->| (Centralized Ingest)  |
|                   |          |                       |
|                   |          | Observability Instance|
|                   |          | (Loki/Grafana)        |
+-------------------+          +-----------------------+
✅ 📐 Example Consistency
🤖 STACKIT Model Serving

✅ Example follows repository conventions.

🔍 STACKIT Cloud Advisor

Repository Convention Review

The provided git diff has been reviewed against the established repository conventions and STACKIT-specific provider requirements. While the example is structurally sound and follows most naming and documentation standards, there are specific deviations regarding provider versioning and variable documentation.

❌ Deviations from Conventions

1. Missing Variable Descriptions
The convention requires all variables to have a description attribute. In 020-variables.tf, the stackit_project_id variable is missing this attribute.

variable "stackit_project_id" {
  type        = string
  description = "The STACKIT Project ID where resources will be provisioned."
  default     = "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}

2. Non-Explicit Provider Version Constraints
The convention requires all providers in required_providers blocks to have an explicit version constraint. In 010-provider.tf, the versions use the >= operator, which allows for breaking changes in minor or major updates. For stable repository examples, a pessimistic constraint (e.g., ~>) or a fixed version is preferred to ensure reproducibility.

terraform {
  required_version = ">= 1.11.0"
  required_providers {
    stackit = {
      source  = "stackitcloud/stackit"
      version = "~> 0.113.0"
    }
    kubernetes = {
      source  = "hashicorp/kubernetes"
      version = "~> 2.30.0"
    }
  }
}

🔍 STACKIT-Specific Observations

  • Beta Resource Usage: The example correctly utilizes enable_beta_resources = true in the stackit provider block. This is a requirement for accessing certain experimental or preview features, such as the SKE audit logging mentioned in the README Stackit Terraform Provider.
  • Experiments Flag: The use of experiments = ["ske"] is correctly implemented to support the ephemeral resource stackit_ske_kubeconfig.
  • Authentication Flow: The example uses the Key Flow via service_account_key_path, which is a standard and supported method for authenticating the STACKIT Terraform Provider Docs.
  • Topology Note: The architecture correctly implements a "Telemetry Link" to connect the project stream to the Telemetry Router, which is the required mechanism for forwarding logs to an Observability instance Docs.

Summary Table

Feature Status Note
3-digit Prefixes ✅ Pass Files follow 010-, 020-, etc.
README/MAINTAINERS ✅ Pass Both files are present and correctly formatted.
Variable Naming ✅ Pass All variables use snake_case.
Variable Descriptions ❌ Fail stackit_project_id is missing a description.
Provider Constraints ❌ Fail Uses >= instead of explicit/pessimistic constraints.
License Headers ✅ Pass Apache 2.0 headers are present on all .tf files.
Lock File ✅ Pass .terraform.lock.hcl is present and committed.
✅ 📚 Example README
🤖 STACKIT Model Serving

✅ Example READMEs are complete.

🔍 STACKIT Cloud Advisor

Review of examples/ske-kubeapi-audit-log/

I have reviewed the provided git diff for the new example directory. Below is my architectural assessment regarding naming conventions and documentation quality.

1. Naming Convention Analysis

The directory name ske-kubeapi-audit-log is compliant with our standards.

  • Clarity: It explicitly identifies the primary STACKIT service involved (SKE - STACKIT Kubernetes Engine).
  • Accuracy: It accurately describes the specific use-case (capturing and processing kube-apiserver audit logs).
  • Consistency: The name follows the established pattern of [service]-[use-case], making it easy for users to locate via CLI or file explorers.

2. README Quality Assessment

The README.md is high quality and meets all architectural requirements for a production-ready example.

  • Service Identification: It clearly outlines the interplay between several STACKIT services:
    • SKE (with the audit feature enabled).
    • Telemetry Router (acting as the ingestion engine).
    • Telemetry Link (for project-level stream subscription).
    • Observability Instance (providing the managed Loki/Grafana backend).
  • Demonstration Depth: The README goes beyond a simple "how-to" by explaining the Layered Filtering Architecture, which is crucial for an architect to understand how to manage log volume and costs.
  • Usage Section: A clear How to use section is provided, including the necessary terraform init and terraform apply commands.
  • Value Add: The inclusion of a "What gets created" table and specific LogQL queries for filtering "Human activity" vs "System noise" provides immediate practical value to the user.

Architectural Topology

For context, the example demonstrates the following data flow:

[ SKE Cluster ] --(Audit Logs)--> [ Telemetry Link ]
                                        |
                                        v
[ Observability ] <--(OTLP)--- [ Telemetry Router ]
(Loki/Grafana)
Component Role Key Configuration
SKE Log Source audit = { enabled = true }
Telemetry Link Connector Connects Project ID to Router
Telemetry Router Processor Applies filter (e.g., kubernetes-audit)
Observability Sink Managed Loki/Grafana endpoint

Verdict:
✅ Example READMEs are complete.

✅ 💬 Commit Messages
🤖 STACKIT Model Serving

✅ Commit messages are descriptive.

🔍 STACKIT Cloud Advisor

Based on the criteria provided, I have evaluated the commit message from your pull request.

Commit Message Evaluation

The provided commit message is:
example(ske): add example for kubeapi audit logs

Status Message Analysis
✅ Pass example(ske): add example for kubeapi audit logs This message is sufficiently descriptive. It follows a structured format (likely Conventional Commits), identifies the scope/context (ske), and clearly describes the action taken (adding an example) and the specific resource being addressed (kubeapi audit logs).

Architect's Summary

All messages in the provided list are sufficiently descriptive. No vague or contextless messages (such as 'fix', 'update', or 'wip') were detected in this specific sample.

✅ Commit messages are descriptive.

⚠️ 🏷️ Tag Quality
🤖 STACKIT Model Serving
  • The example README.md is missing the required <!-- tags: ... --> comment on line 1.
  • The tags should include: ske, telemetry-router, observability, audit-logs, kubernetes, loki, grafana, filtering, otlp
  • Suggested fix: Add this line at the top of examples/ske-kubeapi-audit-log/README.md:
🔍 STACKIT Cloud Advisor

I have reviewed the provided git diff for the new example examples/ske-kubeapi-audit-log/.

While the Terraform code correctly implements the architecture for shipping SKE audit logs via a Telemetry Router to an Observability instance, the documentation (README.md) is missing the required metadata tags.

🚩 Missing Tags

The new example examples/ske-kubeapi-audit-log/README.md does not contain the mandatory <!-- tags: ... --> line on line 1. Based on the resources implemented in the .tf files, the tags must include the relevant STACKIT products and the open-source patterns used.

Analysis of resources touched:

  • stackit_ske_cluster \rightarrow ske
  • stackit_observability_instance \rightarrow observability
  • stackit_telemetryrouter_instance \rightarrow telemetry-router
  • stackit_network \rightarrow iaas
  • Pattern: Uses OpenTelemetry for log ingestion.

🛠️ Suggested Correction

Please add the following tag line to the very first line of examples/ske-kubeapi-audit-log/README.md:

<!-- tags: ske, observability, telemetry-router, iaas, opentelemetry -->

Summary of Review

Resource Path Status Issue Suggestion
examples/ske-kubeapi-audit-log/README.md ❌ Failed Missing tag line on line 1. Add <!-- tags: ske, observability, telemetry-router, iaas, opentelemetry -->

Generated automatically — treat as a hint, not a gate.

## 🤖 AI PR Review > [`c103c361`](https://professional-service.git.onstackit.cloud/professional-service-best-practices/professional-service/commit/c103c3615151761f24b531d0005aaa41873f1b55) · STACKIT Model Serving & STACKIT Cloud Advisor <details> <summary>⚠️ 📝 Spelling & Grammar</summary> <details> <summary>🤖 STACKIT Model Serving</summary> - `examples/ske-kubeapi-audit-log/README.md:12` — “comming” → “coming” - `examples/ske-kubeapi-audit-log/README.md:13` — “overwhleming” → “overwhelming” - `examples/ske-kubeapi-audit-log/README.md:20` — “review” → “preview” - `examples/ske-kubeapi-audit-log/README.md:75` — “have a managed Loki” → “has a managed Loki” </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### Spelling and Grammar Review I have reviewed the provided git diff for spelling and grammar errors in the prose content, including Markdown documentation, Terraform variable descriptions, and inline comments. I have also verified the usage of STACKIT product names and technical terminology against the provided documentation. The following errors were identified: **File: `examples/ske-kubeapi-audit-log/010-provider.tf`** * **Line 18:** `distributed under the License is distributed on an "AS IS" BASIS,` $\rightarrow$ `distributed under the License is distributed on an "AS IS" basis,` *(Note: While this is part of a standard Apache License header, "BASIS" is capitalized in a way that deviates from standard prose, though often accepted in legal headers. If strictly following grammar rules for prose, it should be lowercase.)* **File: `examples/ske-kubeapi-audit-log/README.md`** * **Line 10:** `doing `get`/`list`/`watch` calls comming from kublet or gardener` $\rightarrow$ `doing `get`/`list`/`watch` calls coming from kubelet or gardener` * **Line 10:** `This can be overwhleming so this example aims to show how you can view relevant audit logs as well.` $\rightarrow$ `This can be overwhelming, so this example aims to show how you can view relevant audit logs as well.` * **Line 16:** `because this feature is in private review` $\rightarrow$ `because this feature is in private preview` * **Line 54:** `Both are Loki underneath. STACKIT Logs only focusing on have a managed Loki and Observability comes with a managed Grafana on top.` $\rightarrow$ `Both are Loki underneath. STACKIT Logs only focuses on having a managed Loki, and Observability comes with a managed Grafana on top.` * **Line 138:** `The apply already produces findable audit records: the canary Namespace and ConfigMap are created` $\rightarrow$ `The apply already produces findable audit records: the canary namespace and ConfigMap are created.` *** **Architect's Note on Terminology:** The usage of **SKE** (STACKIT Kubernetes Engine), **Telemetry Router**, **Observability**, and **Telemetry Link** is consistent with STACKIT product naming conventions. The technical distinction between **STACKIT Logs** and **STACKIT Observability** is correctly maintained in the prose. </details> </details> <details> <summary>⚠️ 🏗️ Infrastructure Changes</summary> <details> <summary>🤖 STACKIT Model Serving</summary> - Creates STACKIT network for SKE nodes with IPv4 prefix and nameservers - Creates STACKIT Observability instance (Loki + Grafana) with log retention and ACL - Creates STACKIT Telemetry Router instance and associated access token - Creates STACKIT Telemetry Link to route project logs to the router - Creates STACKIT SKE Kubernetes cluster with audit enabled, node pools, and network attachment - Creates ephemeral kubeconfig for cluster access during Terraform apply - Deploys Kubernetes namespace and config map via ephemeral kubeconfig as audit canary - Exposes outputs for cluster name, router ID/URI, link ID, Grafana URL, and kubeconfig command ⚠️ No destructive changes detected — all resources are newly created. No force-replacements or deletions in this diff. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### 1. Identified STACKIT Services Based on the Terraform configuration, the following services are being provisioned or utilized: * **SKE (STACKIT Kubernetes Engine):** A cluster is being provisioned (`stackit_ske_cluster`) with audit logging enabled. It uses a custom network and a specific node pool configuration. * **Network:** A dedicated STACKIT network (`stackit_network`) is being created to host the SKE nodes. * **Observability:** A managed instance (`stackit_observability_instance`) is being provisioned using the `Observability-Large-EU01` plan to store and visualize logs. * **Telemetry Router:** A managed router instance (`stackit_telemetryrouter_instance`) is being deployed to act as the ingestion point for audit logs. * **Telemetry Link:** A link (`stackit_telemetrylink`) is being established to connect the STACKIT project to the Telemetry Router. **Data Flow Architecture:** ```text [ SKE Cluster ] --(Audit Logs)--> [ Telemetry Link ] | v [ Observability ] <--(OTLP/Auth)-- [ Telemetry Router ] ``` --- ### 2. STACKIT Best Practices & Architectural Review #### **Security & Access Control** * **Telemetry ACL:** The variable `telemetry_acl` is defaulted to `["0.0.0.0/0"]`. **Recommendation:** For production environments, restrict this to specific known IP ranges to minimize the attack surface of your Observability instance. * **Router Placement:** The code includes a comment noting that for demo purposes, the router lives in the same project as the cluster. **Best Practice:** In a production-grade architecture, you should deploy the Telemetry Router in a dedicated "Hub" project to centralize audit log collection from multiple "Spoke" projects [Architecture](https://docs.stackit.cloud/de/products/logging-and-monitoring/telemetry-router/basics/architecture/). * **Credential Management:** The use of `stackit_observability_credential` to pass technical credentials to the router is the correct way to handle OTLP authentication via Basic Auth. #### **Observability Configuration** * **Log Retention:** The `observability_logs_retention_days` is set to `7`. * **Note:** The `Observability-Large` plan supports a default of **5 days** and can be configured up to **30 days** via the API [Service Plans Observability](https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/). Ensure your retention requirement aligns with your compliance needs. * **Filtering:** The `stackit_telemetryrouter_destination` uses a specific filter for `stackit.log.kind = "kubernetes-audit"`. This is an excellent practice to prevent "noise" and unnecessary costs by only routing relevant audit data to the Observability instance. --- ### 3. Quotas, Regional Constraints, and Limitations #### **Regional Constraints** * **Telemetry Router Availability:** The Telemetry Router is currently only available in the **EU01 (Germany)** and **EU02 (Austria)** regions [Scope and Limitations](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/basics/overview/). * **Telemetry Link Locality:** A critical constraint is that the **Telemetry Link requires the Telemetry Router to be in the same region** as the resources being linked [Architecture](https://docs.stackit.cloud/de/products/logging-and-monitoring/telemetry-router/basics/architecture/). Since your `stackit_region` defaults to `eu01`, ensure your cluster and router remain in this region. #### **Service Limitations** * **Telemetry Router Data Type:** The router is currently specialized for **Audit Logs only**. You cannot use this specific router instance to send custom application metrics or traces [Scope and Limitations](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/basics/overview/). * **Observability Plan Limits:** You are using the `Observability-Large` plan. Be aware of the following hard limits for this plan: * **Logs Size:** Up to **1,000 GB** [Service Plans Observability](https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/). * **Logs Retention:** Maximum of **30 days** [Service Plans Observability](https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/). * **Metrics Samples:** Up to **300,000** samples per minute [Service Plans Observability](https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/). * **Observability Backup:** Note that while configuration and dashboards are backed up automatically, there is **no backup of the telemetry data itself** (metrics, logs, or traces) [Service Certificate](https://docs.stackit.com/en/asset/download/34321/file/Service_certificate_STACKIT_Observability.pdf?version=10). </details> </details> <details> <summary>⚠️ 🔒 Security Review</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ No security issues found. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### 1. STACKIT IAM or Authorization Misconfigurations * 🟡 **Medium: Telemetry Router Access Token Scope** The `stackit_telemetryrouter_access_token` is created to allow the `stackit_telemetrylink` to subscribe the project to the router. While this is the functional requirement for the link, ensure that the principle of least privilege is applied to the identity using this token if it were to be exported or used outside of the automated Terraform flow. ### 2. Missing STACKIT-Specific Security Controls * 🔴 **High: Overly Permissive Observability ACL** In `020-variables.tf`, the `telemetry_acl` variable defaults to `["0.0.0.0/0"]`. The **Observability** service is available on the Internet by default [How To Control Instance Access](https://docs.stackit.cloud/products/logging-and-monitoring/observability/how-tos/how-to-control-instance-access/). Using `0.0.0.0/0` allows full access to Grafana, Metrics, Logs, and Traces from any IP address [How To Control Instance Access](https://docs.stackit.cloud/products/logging-and-monitoring/observability/how-tos/how-to-control-instance-access/). * **Recommendation:** Restrict this to specific corporate IP ranges or use the `stackit_public_ip_ranges` data source to at least limit access to known STACKIT service ranges [Public Ip Ranges](https://registry.terraform.io/providers/stackitcloud/stackit/latest/docs/data-sources/public_ip_ranges). * 🔴 **High: Missing SKE API Server Access Control (ACL)** The `stackit_ske_cluster` resource in `070-ske-cluster.tf` does not define the `extensions.acl` block. Without this, the Kubernetes API server is not protected by the SKE ACL functionality, which is a critical layer of security for limiting access to the API [Enhance The Security Of Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/enhance-the-security-of-your-cluster/) [Configure Access Control For Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/configure-access-control-for-your-cluster/). * **Recommendation:** Implement the `acl` extension within the SKE cluster configuration to restrict `allowedCidrs` to trusted networks [Configure Access Control For Your Cluster](https://docs.stackit.cloud/de/products/runtime/kubernetes-engine/how-tos/configure-access-control-for-your-cluster/) [Configure Access Control For Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/configure-access-control-for-your-cluster/). * 🟡 **Medium: Lack of Network Policies** The current configuration defines a `stackit_network` for the nodes, but there are no Kubernetes `NetworkPolicy` resources defined to control pod-to-pod traffic [Enhance The Security Of Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/enhance-the-security-of-your-cluster/). * **Recommendation:** Implement a "Default Deny" policy for the `audit-canary` namespace and subsequent production namespaces to enforce granular traffic control [Enhance The Security Of Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/enhance-the-security-of-your-cluster/). * 🟡 **Medium: Absence of Pod Security Standards (PSS)** The `audit-canary` namespace is created without any Pod Security Standard labels [Enhance The Security Of Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/enhance-the-security-of-your-cluster/). * **Recommendation:** Apply `pod-security.kubernetes.io/enforce: restricted` labels to namespaces to prevent privileged containers and enforce non-root users [Enhance The Security Of Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/enhance-the-security-of-your-cluster/). ### 3. Secrets and Sensitive Values * 🟢 **Low: Service Account Key Path** The variable `stackit_service_account_key_path` is used to point to a JSON key file. While the path itself is not a secret, the security of the entire deployment relies on the protection of the file at that path. Ensure this file is never committed to version control. ### 4. STACKIT Compliance and Audit-Logging Considerations * 🟢 **Low: Audit Log Retention** The `observability_logs_retention_days` is set to `7` days [020-variables.tf]. While this is sufficient for a "demo" (as suggested by the code comments), production compliance requirements often mandate longer retention periods for audit logs. * 🟢 **Low: Telemetry Router Architecture** The code notes that the router lives in the same project as the cluster for demo purposes [050-telemetry-router.tf]. For production environments, a more secure architecture involves running the **Telemetry Router** in a dedicated "hub" project to centralize logging and isolate the observability infrastructure from workload projects. ```ascii [ Project: Workload ] [ Project: Hub/Logging ] +-------------------+ +-----------------------+ | SKE Cluster | | Telemetry Router | | (Audit Logs) ---->|--------->| (Centralized Ingest) | | | | | | | | Observability Instance| | | | (Loki/Grafana) | +-------------------+ +-----------------------+ ``` </details> </details> <details> <summary>✅ 📐 Example Consistency</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ Example follows repository conventions. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### Repository Convention Review The provided git diff has been reviewed against the established repository conventions and STACKIT-specific provider requirements. While the example is structurally sound and follows most naming and documentation standards, there are specific deviations regarding provider versioning and variable documentation. #### ❌ Deviations from Conventions **1. Missing Variable Descriptions** The convention requires *all* variables to have a `description` attribute. In `020-variables.tf`, the `stackit_project_id` variable is missing this attribute. ```hcl variable "stackit_project_id" { type = string description = "The STACKIT Project ID where resources will be provisioned." default = "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx" } ``` **2. Non-Explicit Provider Version Constraints** The convention requires all providers in `required_providers` blocks to have an *explicit* version constraint. In `010-provider.tf`, the versions use the `>=` operator, which allows for breaking changes in minor or major updates. For stable repository examples, a pessimistic constraint (e.g., `~>`) or a fixed version is preferred to ensure reproducibility. ```hcl terraform { required_version = ">= 1.11.0" required_providers { stackit = { source = "stackitcloud/stackit" version = "~> 0.113.0" } kubernetes = { source = "hashicorp/kubernetes" version = "~> 2.30.0" } } } ``` #### 🔍 STACKIT-Specific Observations * **Beta Resource Usage:** The example correctly utilizes `enable_beta_resources = true` in the `stackit` provider block. This is a requirement for accessing certain experimental or preview features, such as the SKE audit logging mentioned in the README [Stackit Terraform Provider](https://docs.stackit.cloud/developer-tools/stackit-iac/stackit-terraform-provider/). * **Experiments Flag:** The use of `experiments = ["ske"]` is correctly implemented to support the `ephemeral` resource `stackit_ske_kubeconfig`. * **Authentication Flow:** The example uses the **Key Flow** via `service_account_key_path`, which is a standard and supported method for authenticating the STACKIT Terraform Provider [Docs](https://registry.terraform.io/providers/stackitcloud/stackit/latest/docs). * **Topology Note:** The architecture correctly implements a "Telemetry Link" to connect the project stream to the Telemetry Router, which is the required mechanism for forwarding logs to an **Observability instance** [Docs](https://registry.terraform.io/providers/stackitcloud/stackit/latest/docs). #### Summary Table | Feature | Status | Note | | :--- | :--- | :--- | | **3-digit Prefixes** | ✅ Pass | Files follow `010-`, `020-`, etc. | | **README/MAINTAINERS** | ✅ Pass | Both files are present and correctly formatted. | | **Variable Naming** | ✅ Pass | All variables use `snake_case`. | | **Variable Descriptions** | ❌ Fail | `stackit_project_id` is missing a description. | | **Provider Constraints** | ❌ Fail | Uses `>=` instead of explicit/pessimistic constraints. | | **License Headers** | ✅ Pass | Apache 2.0 headers are present on all `.tf` files. | | **Lock File** | ✅ Pass | `.terraform.lock.hcl` is present and committed. | </details> </details> <details> <summary>✅ 📚 Example README</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ Example READMEs are complete. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### Review of `examples/ske-kubeapi-audit-log/` I have reviewed the provided git diff for the new example directory. Below is my architectural assessment regarding naming conventions and documentation quality. #### 1. Naming Convention Analysis The directory name `ske-kubeapi-audit-log` is **compliant** with our standards. * **Clarity:** It explicitly identifies the primary STACKIT service involved (**SKE** - STACKIT Kubernetes Engine). * **Accuracy:** It accurately describes the specific use-case (capturing and processing **kube-apiserver audit logs**). * **Consistency:** The name follows the established pattern of `[service]-[use-case]`, making it easy for users to locate via CLI or file explorers. #### 2. README Quality Assessment The `README.md` is **high quality** and meets all architectural requirements for a production-ready example. * **Service Identification:** It clearly outlines the interplay between several STACKIT services: * **SKE** (with the `audit` feature enabled). * **Telemetry Router** (acting as the ingestion engine). * **Telemetry Link** (for project-level stream subscription). * **Observability Instance** (providing the managed Loki/Grafana backend). * **Demonstration Depth:** The README goes beyond a simple "how-to" by explaining the **Layered Filtering Architecture**, which is crucial for an architect to understand how to manage log volume and costs. * **Usage Section:** A clear `How to use` section is provided, including the necessary `terraform init` and `terraform apply` commands. * **Value Add:** The inclusion of a "What gets created" table and specific **LogQL queries** for filtering "Human activity" vs "System noise" provides immediate practical value to the user. #### Architectural Topology For context, the example demonstrates the following data flow: ```text [ SKE Cluster ] --(Audit Logs)--> [ Telemetry Link ] | v [ Observability ] <--(OTLP)--- [ Telemetry Router ] (Loki/Grafana) ``` | Component | Role | Key Configuration | | :--- | :--- | :--- | | **SKE** | Log Source | `audit = { enabled = true }` | | **Telemetry Link** | Connector | Connects Project ID to Router | | **Telemetry Router** | Processor | Applies `filter` (e.g., `kubernetes-audit`) | | **Observability** | Sink | Managed Loki/Grafana endpoint | **Verdict:** ✅ Example READMEs are complete. </details> </details> <details> <summary>✅ 💬 Commit Messages</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ Commit messages are descriptive. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> Based on the criteria provided, I have evaluated the commit message from your pull request. ### Commit Message Evaluation The provided commit message is: `example(ske): add example for kubeapi audit logs` | Status | Message | Analysis | | :--- | :--- | :--- | | ✅ **Pass** | `example(ske): add example for kubeapi audit logs` | This message is sufficiently descriptive. It follows a structured format (likely Conventional Commits), identifies the scope/context (`ske`), and clearly describes the action taken (adding an example) and the specific resource being addressed (`kubeapi audit logs`). | ### Architect's Summary All messages in the provided list are sufficiently descriptive. No vague or contextless messages (such as 'fix', 'update', or 'wip') were detected in this specific sample. ✅ Commit messages are descriptive. </details> </details> <details> <summary>⚠️ 🏷️ Tag Quality</summary> <details> <summary>🤖 STACKIT Model Serving</summary> - The example README.md is missing the required `<!-- tags: ... -->` comment on line 1. - The tags should include: ske, telemetry-router, observability, audit-logs, kubernetes, loki, grafana, filtering, otlp - Suggested fix: Add this line at the top of examples/ske-kubeapi-audit-log/README.md: <!-- tags: ske, telemetry-router, observability, audit-logs, kubernetes, loki, grafana, filtering, otlp --> </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> I have reviewed the provided git diff for the new example `examples/ske-kubeapi-audit-log/`. While the Terraform code correctly implements the architecture for shipping SKE audit logs via a **Telemetry Router** to an **Observability** instance, the documentation (README.md) is missing the required metadata tags. ### 🚩 Missing Tags The new example `examples/ske-kubeapi-audit-log/README.md` does not contain the mandatory `<!-- tags: ... -->` line on line 1. Based on the resources implemented in the `.tf` files, the tags must include the relevant STACKIT products and the open-source patterns used. **Analysis of resources touched:** * `stackit_ske_cluster` $\rightarrow$ `ske` * `stackit_observability_instance` $\rightarrow$ `observability` * `stackit_telemetryrouter_instance` $\rightarrow$ `telemetry-router` * `stackit_network` $\rightarrow$ `iaas` * Pattern: Uses **OpenTelemetry** for log ingestion. ### 🛠️ Suggested Correction Please add the following tag line to the very first line of `examples/ske-kubeapi-audit-log/README.md`: ```markdown <!-- tags: ske, observability, telemetry-router, iaas, opentelemetry --> ``` ### Summary of Review | Resource Path | Status | Issue | Suggestion | | :--- | :--- | :--- | :--- | | `examples/ske-kubeapi-audit-log/README.md` | ❌ **Failed** | Missing tag line on line 1. | Add `<!-- tags: ske, observability, telemetry-router, iaas, opentelemetry -->` | </details> </details> --- _Generated automatically — treat as a hint, not a gate._
marc.adrian force-pushed example/ske-kubeapi-audit-log from c103c36151
Some checks failed
Default CI / Check for Open TODOs (pull_request) Successful in 43s
Default CI / Check README Tags (pull_request) Failing after 1m5s
Default CI / Check AGENTS.md is up to date (pull_request) Has been skipped
AI PR Review / AI PR Review (pull_request) Successful in 2m7s
Default CI / Pre-Commit Hooks (pull_request) Failing after 2m14s
to c4fc7ef1bf
Some checks failed
Default CI / Check for Open TODOs (pull_request) Successful in 1m30s
Default CI / Check README Tags (pull_request) Successful in 1m25s
AI PR Review / AI PR Review (pull_request) Successful in 2m51s
Default CI / Check AGENTS.md is up to date (pull_request) Failing after 45s
Default CI / Pre-Commit Hooks (pull_request) Failing after 3m12s
2026-09-22 08:07:08 +00:00
Compare

🤖 AI PR Review

c4fc7ef1 · STACKIT Model Serving & STACKIT Cloud Advisor

⚠️ 📝 Spelling & Grammar
🤖 STACKIT Model Serving
  • AGENTS.md:150: "This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the STACKIT Telemetry Router into an Observability instance, where they can be analysed with LogQL in Grafana" → "This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the STACKIT Telemetry Router into an Observability instance, where they can be analyzed with LogQL in Grafana"
  • examples/ske-kubeapi-audit-log/README.md:15: "A lot of logs are produced by the cluster itself, doing get/list/watch calls coming from kublet or gardener (infrastructure to manage your SKE). This can be overwhelming so this example aims to show how you can view relevant audit logs as well." → "A lot of logs are produced by the cluster itself, doing get/list/watch calls coming from kubelet or gardener (infrastructure to manage your SKE). This can be overwhelming, so this example aims to show how you can view relevant audit logs as well."
  • examples/ske-kubeapi-audit-log/README.md:20: "This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the STACKIT Telemetry Router into an Observability instance, where they can be analysed with LogQL in Grafana." → "This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the STACKIT Telemetry Router into an Observability instance, where they can be analyzed with LogQL in Grafana."
  • examples/ske-kubeapi-audit-log/README.md:30: "Please get in contact with STACKIT if you want to enable it." → "Please get in contact with STACKIT if you want to enable it." (no change needed — already correct)
  • examples/ske-kubeapi-audit-log/README.md:37: "This example only takes the project ID; it does not create the organization, folder, project or network area." → "This example only takes the project ID; it does not create the organization, folder, project, or network area."
  • examples/ske-kubeapi-audit-log/README.md:45: "the SKE audit logs feature is enabled for your account, because this feature is in private review" → "the SKE audit logs feature is enabled for your account, because this feature is in private preview"
  • examples/ske-kubeapi-audit-log/README.md:83: "STACKIT Logs focuses only on providing a managed Loki and Observability includes a managed Grafana on top." → "STACKIT Logs focuses only on providing a managed Loki, and Observability includes a managed Grafana on top."
  • examples/ske-kubeapi-audit-log/README.md:91: "In this case it's easier for testing purposes." → "In this case, it’s easier for testing purposes."
  • examples/ske-kubeapi-audit-log/README.md:101: "The audit policy is managed by SKE and cannot be tuned in any way, so you cannot tell the kube-apiserver to skip watch calls from system:serviceaccount:kube-system:* for example." → "The audit policy is managed by SKE and cannot be tuned in any way, so you cannot tell the kube-apiserver to skip watch calls from system:serviceaccount:kube-system:*, for example."
  • examples/ske-kubeapi-audit-log/README.md:113: "In this case we leave it open." → "In this case, we leave it open."
  • examples/ske-kubeapi-audit-log/README.md:127: "STACKIT sets on every forwarded record (dots become underscores when they arrive as Loki labels):" → "STACKIT sets on every forwarded record (dots become underscores when they arrive as Loki labels):" (no change needed — already correct)
  • examples/ske-kubeapi-audit-log/README.md:167: "Open the grafana_url output. Get credentials via portal." → "Open the grafana_url output. Get credentials via the portal."
  • examples/ske-kubeapi-audit-log/README.md:174: "logql\n{service_name=\"ske\", service_instance_id=\"audit-demo\"}\n" → "logql\n{service_name=\"ske\", service_instance_id=\"audit-demo\"}\n" (no change needed — already correct)
  • examples/ske-kubeapi-audit-log/README.md:180: "```logql\n{service_name="ske", stackit_log_kind="kubernetes-audit", service_instance_id="audit-demo
🔍 STACKIT Cloud Advisor

✅ No spelling or grammar issues found.

⚠️ 🏗️ Infrastructure Changes
🤖 STACKIT Model Serving
  • Creates STACKIT SKE cluster with audit enabled and connected to Telemetry Router
  • Creates STACKIT network for SKE nodes with specified IPv4 prefix and nameservers
  • Creates STACKIT Observability instance for audit log storage and Grafana access
  • Creates STACKIT Telemetry Router with access token and destination for Kubernetes audit logs
  • Creates STACKIT Telemetry Link to route project logs to the Telemetry Router
  • Deploys Kubernetes namespace and ConfigMap via ephemeral kubeconfig for audit log generation
  • Exposes outputs for cluster name, router ID/URI, Grafana URL, and kubeconfig command

⚠️ No destructive changes detected. All resources are new creations.

🔍 STACKIT Cloud Advisor

1. Identified STACKIT Services

The pull request provisions a complete telemetry pipeline to capture and store SKE audit logs. The following services are being introduced:

  • SKE (STACKIT Kubernetes Engine): A new cluster is being provisioned (stackit_ske_cluster.this) with audit logging enabled. It includes a node pool using the Flatcar OS.
  • Network: A dedicated STACKIT network (stackit_network.ske_nodes) is created to host the SKE nodes.
  • Observability: A managed instance (stackit_observability_instance.audit) is provisioned to act as the long-term storage and visualization layer (Loki/Grafana). It includes specific credentials for ingestion.
  • Telemetry Router: A managed router instance (stackit_telemetryrouter_instance.this) is provisioned to act as the central ingestion and distribution point.
  • Telemetry Link: A project-level link (stackit_telemetrylink.this) is established to connect the project's audit stream to the Telemetry Router.

Data Flow Topology:

[ SKE Cluster ] --(Audit Logs)--> [ Telemetry Link ]
                                         |
                                         v
[ Observability ] <--(OTLP)--- [ Telemetry Router ]
(Loki/Grafana)

2. STACKIT Best Practices & Architectural Review

While the implementation is functional, there are several architectural considerations regarding STACKIT best practices:

  • Centralized vs. Localized Routing: The code currently deploys the Telemetry Router within the same project as the SKE cluster. As noted in the documentation, for production environments, it is highly recommended to run a single, central Telemetry Router in a dedicated "hub" project and feed it via Telemetry Links from various workload projects “We recommend a single, central Telemetry Router in a core project...”.
  • Observability Plan Selection: The code uses the Observability-Large-EU01 plan. Ensure this aligns with your expected log volume. Note that the Observability Large plan has a specific limit of 1,000 GB for log storage “Detailed limits for Observability plans”.
  • Security (ACLs): The telemetry_acl is currently set to ["0.0.0.0/0"]. This is highly permissive. For a production environment, you should restrict access to the Observability instance to known IP ranges or specific internal networks.
  • Observability Ingestion Path: The stackit_telemetryrouter_destination uses the otlp_http_logs_url from the Observability instance. This is correct, but ensure that the destination URI in any manual configurations includes the full OTLP-compliant path ending in /v1/logs, as the router does not append this automatically “The destination URI must include the full OTLP-compliant path for logs, ending in /v1/logs”.

3. Service Quotas, Constraints, and Limitations

Before applying this change, please be aware of the following technical constraints:

Constraint Category Detail Impact/Requirement
Telemetry Router Scope Only handles Audit Logs “Currently, the service exclusively handles Audit Logs.” You cannot use this specific router instance to send application logs or custom metrics.
Telemetry Link Uniqueness Only one Link per hierarchy object “Each hierarchy object can only be connected to one Telemetry Router at a time.” If this project is already linked to another router, this deployment will fail.
Regional Constraint Links must target a Router in the same region “A Telemetry Link can only target a Router in the same region.” Ensure var.stackit_region is consistent across the SKE cluster, the Router, and the Observability instance.
Observability Retention Default log retention is 5 days “Default retention: 5 days (up to 30 days)” The variable observability_logs_retention_days is set to 7. This is within the allowed range (up to 30 days), but ensure it meets your compliance needs.
Observability Quotas Storage and sample limits apply “The number of metrics samples, alerting rules and data storage capacity increase with larger plan sizes.” Monitor your log storage usage. It is recommended to set up alerts when consumption reaches 90% of your plan's limit “It is highly recommended to set up early warning alerts... when your resource consumption reaches 90%”.
⚠️ 🔒 Security Review
🤖 STACKIT Model Serving

✅ No security issues found.

🔍 STACKIT Cloud Advisor

Security Review Findings

Based on the provided git diff, I have identified several security considerations regarding the deployment of the SKE cluster, the Observability instance, and the Telemetry Router.

1. STACKIT IAM & Authorization Misconfigurations

  • 🔴 High: Overly Permissive Observability Access Control (ACL)
    The variable telemetry_acl in 020-variables.tf is defaulted to ["0.0.0.0/0"]. This is applied to the stackit_observability_instance in 040-observability.tf. In a production environment, allowing open access to your observability/logging endpoint is a significant risk. It is recommended to restrict this to known corporate IP ranges or specific management networks Enhance The Security Of Your Cluster Enhance The Security Of Your Cluster.
  • 🟡 Medium: Telemetry Router Deployment Context
    The code notes that for demo purposes, the router lives in the same project as the cluster [050-telemetry-router.tf]. While acceptable for a proof-of-concept, a production architecture should ideally run the Telemetry Router in a dedicated "hub" project to isolate the telemetry ingestion plane from the workload plane.

2. Missing STACKIT-Specific Security Controls

  • 🔴 High: Missing SKE API Server Access Control (ACL)
    The stackit_ske_cluster resource in 070-ske-cluster.tf does not implement the extensions.acl block. Without this, the Kubernetes API server is not restricted by source IP ranges. For production, you should implement the acl extension to limit access to trusted CIDRs, such as corporate VPNs or CI/CD egress IPs Configure Access Control For Your Cluster.
  • 🟡 Medium: Lack of Network Policy Baseline
    The pull request defines a stackit_network for SKE nodes, but there are no Kubernetes NetworkPolicy resources defined to control pod-to-pod traffic. By default, pods in a cluster can communicate broadly Containers & Kubernetes Security — Network Security & Cluster Hardening. I recommend implementing a "default-deny" ingress policy as a starting point for all namespaces Enhance The Security Of Your Cluster Containers & Kubernetes Security — Network Security & Cluster Hardening.
  • 🟡 Medium: Absence of Private Cluster Configuration
    The current configuration does not explicitly enable private cluster features. Depending on your security requirements, you may want to evaluate if the API server and node communication should be restricted to private networks to reduce the public attack surface Enhance The Security Of Your Cluster.

3. Secrets and Sensitive Values

  • 🟢 Low: Service Account Key Path
    The provider configuration in 010-provider.tf uses var.stackit_service_account_key_path. As long as this variable points to a local file path on the runner and is not hard-coded as a string within the .tf files, this follows standard practice. However, ensure the JSON key itself is never committed to version control.

4. Compliance and Audit-Logging Considerations

  • 🟢 Low: SKE Audit Logging Enabled
    The configuration correctly enables SKE audit logging via audit = { enabled = true } in 070-ske-cluster.tf. This is a critical requirement for capturing security-relevant events such as Secret access or RBAC changes Containers & Kubernetes Security — Secure Operations on Kubernetes.
  • 🟢 Low: Centralized Log Routing
    The implementation of the stackit_telemetryrouter_instance and stackit_telemetrylink provides a robust mechanism for centralizing SKE kube-apiserver audit logs into the STACKIT Observability instance. This supports the requirement for centralized logging and long-term retention Containers & Kubernetes Security — Secure Operations on Kubernetes.

Summary Table

Finding Severity Category Recommendation
Open Observability ACL (0.0.0.0/0) 🔴 High Authorization Restrict telemetry_acl to trusted CIDRs.
Missing SKE API ACL 🔴 High Network Isolation Implement extensions.acl in stackit_ske_cluster.
Missing Network Policies 🟡 Medium Network Isolation Deploy a default-deny NetworkPolicy baseline.
Single-Project Router 🟡 Medium Architecture Move Telemetry Router to a dedicated hub project for production.
SKE Audit Logging ✅ Pass Compliance Audit logging is correctly enabled.
✅ 📐 Example Consistency
🤖 STACKIT Model Serving

✅ Example follows repository conventions.

🔍 STACKIT Cloud Advisor

The provided git diff has been reviewed against the established repository conventions and STACKIT-specific provider requirements. While the example is largely well-structured, there are several deviations regarding repository conventions and provider configuration.

🔍 Repository Convention Deviations

Convention Status Observation
3-digit numeric prefixes ✅ Pass Files follow the 010-, 020- pattern.
README.md & MAINTAINERS.md ✅ Pass Both files are present.
Variable descriptions & snake_case ❌ Fail stackit_project_id in 020-variables.tf lacks a description attribute.
Explicit provider version constraints ❌ Fail The required_providers block in 010-provider.tf uses >= constraints instead of pinning to an exact version.
.terraform.lock.hcl committed ✅ Pass The lock file is present in the diff.
Apache 2.0 license headers ✅ Pass Headers are present on all .tf files.

Fix: Missing Variable Description

The variable stackit_project_id must include a description to comply with the requirement that all variables have a description attribute.

variable "stackit_project_id" {
  type        = string
  description = "The STACKIT project ID."
  default     = "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}

Fix: Provider Version Pinning

To ensure environment stability and comply with the requirement for explicit version constraints, avoid using range operators like >= in the required_providers block.

terraform {
  required_version = ">= 1.11.0"
  required_providers {
    stackit = {
      source  = "stackitcloud/stackit"
      version = "0.113.0"
    }
    kubernetes = {
      source  = "hashicorp/kubernetes"
      version = "3.2.1"
    }
  }
}

☁️ STACKIT Provider & Usage Observations

1. Authentication Method

The example uses the Key flow via service_account_key_path STACKIT Terraform Provider. This is the recommended method for most projects Terraform Fundamentals — Setting Up the Environment and Authentication. However, for CI/CD workloads, I recommend transitioning to Workload Identity Federation (WIF) to eliminate long-lived secrets STACKIT Terraform Provider.

2. Beta Resources & Experiments

The configuration correctly enables beta resources and specific experiments required for this architecture:

3. S3 Backend Note

While the example does not implement a remote backend, please note that if you move to an S3 backend using STACKIT Object Storage, setting use_lockfile = true currently has no effect STACKIT Terraform Provider. For state locking, a PostgreSQL backend is the recommended alternative STACKIT Terraform Provider.

4. SKE Audit Logs (Private Preview)

As noted in the README.md, the audit = { enabled = true } configuration for stackit_ske_cluster is currently in private preview [README]. Ensure your account is explicitly enabled by STACKIT before attempting to apply this configuration.

✅ 📚 Example README
🤖 STACKIT Model Serving

✅ Example READMEs are complete.

🔍 STACKIT Cloud Advisor

Review of examples/ske-kubeapi-audit-log

I have reviewed the git diff for the new example directory. Below is my architectural assessment regarding the naming conventions and documentation quality.

1. Naming Analysis

The directory name ske-kubeapi-audit-log is highly effective and meets our architectural standards for clarity and accuracy.

  • Service Identification: It explicitly identifies the primary STACKIT service involved (SKE - STACKIT Kubernetes Engine).
  • Use-Case Specificity: It clearly defines the technical objective (kubeapi-audit-log), which distinguishes it from general SKE cluster provisioning or workload deployment examples.
  • Consistency: The name aligns perfectly with the Terraform resources being introduced (e.g., stackit_ske_cluster with audit = { enabled = true }) and the logical flow of the telemetry pipeline described in the README.

2. README Quality Assessment

The README.md is comprehensive and high-quality, providing significant value for a cloud architect or DevOps engineer looking to implement this pattern.

  • Service Explanation: It clearly outlines the interaction between SKE, Telemetry Router, Telemetry Link, and the Observability instance (Loki + Grafana).
  • Technical Depth: It goes beyond a simple "how-to" by explaining the Layered Filtering approach (Source \rightarrow Router Instance \rightarrow Destination \rightarrow Query time), which is critical for managing log volume and costs in a production environment.
  • Usage Section: A clear usage section is provided, including the necessary terraform init and terraform apply commands.
  • Value-Add Content:
    • It includes a detailed table of Loki labels (mapping OpenTelemetry attributes to Loki labels), which is essential for writing effective LogQL queries.
    • It provides concrete LogQL examples for different visibility levels (e.g., filtering out "noise" from gardener/system accounts).
    • It includes a "What gets created" table that maps Terraform files to specific STACKIT resources, aiding in architectural comprehension.

Summary Table

Criteria Status Architect's Notes
Naming ✅ Pass Clear, service-oriented, and accurately describes the audit-log use case.
README Quality ✅ Pass Excellent depth; includes usage, filtering logic, and practical LogQL queries.

✅ Example READMEs are complete.

✅ 💬 Commit Messages
🤖 STACKIT Model Serving

✅ Commit messages are descriptive.

🔍 STACKIT Cloud Advisor

All provided commit messages are sufficiently descriptive.

✅ Commit messages are descriptive.

✅ 🏷️ Tag Quality
🤖 STACKIT Model Serving

✅ README tags accurately reflect the content.

🔍 STACKIT Cloud Advisor

README Tag Review

I have reviewed the README.md tags for the new example examples/ske-kubeapi-audit-log/. While the tags capture the general intent of the repository, they require adjustments to align with STACKIT product naming conventions and to ensure all relevant managed services are represented.

Analysis of examples/ske-kubeapi-audit-log/README.md

The current tag line is:
<!-- tags: ske, kubernetes, audit-log, kube-apiserver, observability, otel, telemetry-router, telemetry-link -->

1. STACKIT Product Naming & Accuracy

2. Open-Source Tools & Patterns

  • otel: While technically correct (referring to OpenTelemetry), the standard convention in this repository for the protocol/pattern is opentelemetry to match the service capabilities Architecture Ship Traces Logs Metrics To Observability Using Opentelemetry.
  • kubernetes / audit-log / kube-apiserver: These are descriptive of the workload/pattern rather than the STACKIT product, which is acceptable for discovery, but they do not conflict with the requirements.

3. Formatting

  • All existing tags follow the lowercase-and-hyphen-separated rule.

Summary of Findings

Resource Current Tag(s) Status Recommendation
SKE Cluster ske ✅ Keep as is.
Observability observability ✅ Keep as is.
Telemetry Router telemetry-router ✅ Keep as is.
Telemetry Link telemetry-link ✅ Keep as is.
OpenTelemetry otel ⚠️ Change to opentelemetry for consistency with documentation Ship Traces Logs Metrics To Observability Using Opentelemetry.
Network None ❌ Add network to reflect the stackit_network resource.

Suggested Correction

To ensure full coverage of the managed resources and consistent naming, please update the first line of examples/ske-kubeapi-audit-log/README.md to:

<!-- tags: ske, kubernetes, audit-log, kube-apiserver, observability, opentelemetry, telemetry-router, telemetry-link, network -->

Generated automatically — treat as a hint, not a gate.

## 🤖 AI PR Review > [`c4fc7ef1`](https://professional-service.git.onstackit.cloud/professional-service-best-practices/professional-service/commit/c4fc7ef1bf99c8060be1309100b2d647429135af) · STACKIT Model Serving & STACKIT Cloud Advisor <details> <summary>⚠️ 📝 Spelling & Grammar</summary> <details> <summary>🤖 STACKIT Model Serving</summary> - AGENTS.md:150: "This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the **STACKIT Telemetry Router** into an **Observability instance**, where they can be analysed with LogQL in Grafana" → "This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the **STACKIT Telemetry Router** into an **Observability instance**, where they can be analyzed with LogQL in Grafana" - examples/ske-kubeapi-audit-log/README.md:15: "A lot of logs are produced by the cluster itself, doing `get`/`list`/`watch` calls coming from kublet or gardener (infrastructure to manage your SKE). This can be overwhelming so this example aims to show how you can view relevant audit logs as well." → "A lot of logs are produced by the cluster itself, doing `get`/`list`/`watch` calls coming from kubelet or gardener (infrastructure to manage your SKE). This can be overwhelming, so this example aims to show how you can view relevant audit logs as well." - examples/ske-kubeapi-audit-log/README.md:20: "This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the **STACKIT Telemetry Router** into an **Observability instance**, where they can be analysed with LogQL in Grafana." → "This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the **STACKIT Telemetry Router** into an **Observability instance**, where they can be analyzed with LogQL in Grafana." - examples/ske-kubeapi-audit-log/README.md:30: "Please get in contact with STACKIT if you want to enable it." → "Please get in contact with STACKIT if you want to enable it." (no change needed — already correct) - examples/ske-kubeapi-audit-log/README.md:37: "This example only takes the project ID; it does not create the organization, folder, project or network area." → "This example only takes the project ID; it does not create the organization, folder, project, or network area." - examples/ske-kubeapi-audit-log/README.md:45: "the SKE audit logs feature is enabled for your account, because this feature is in private review" → "the SKE audit logs feature is enabled for your account, because this feature is in private preview" - examples/ske-kubeapi-audit-log/README.md:83: "STACKIT Logs focuses only on providing a managed Loki and Observability includes a managed Grafana on top." → "STACKIT Logs focuses only on providing a managed Loki, and Observability includes a managed Grafana on top." - examples/ske-kubeapi-audit-log/README.md:91: "In this case it's easier for testing purposes." → "In this case, it’s easier for testing purposes." - examples/ske-kubeapi-audit-log/README.md:101: "The audit policy is managed by SKE and cannot be tuned in any way, so you cannot tell the kube-apiserver to skip `watch` calls from `system:serviceaccount:kube-system:*` for example." → "The audit policy is managed by SKE and cannot be tuned in any way, so you cannot tell the kube-apiserver to skip `watch` calls from `system:serviceaccount:kube-system:*`, for example." - examples/ske-kubeapi-audit-log/README.md:113: "In this case we leave it open." → "In this case, we leave it open." - examples/ske-kubeapi-audit-log/README.md:127: "STACKIT sets on every forwarded record (dots become underscores when they arrive as Loki labels):" → "STACKIT sets on every forwarded record (dots become underscores when they arrive as Loki labels):" (no change needed — already correct) - examples/ske-kubeapi-audit-log/README.md:167: "Open the `grafana_url` output. Get credentials via portal." → "Open the `grafana_url` output. Get credentials via the portal." - examples/ske-kubeapi-audit-log/README.md:174: "```logql\n{service_name=\"ske\", service_instance_id=\"audit-demo\"}\n```" → "```logql\n{service_name=\"ske\", service_instance_id=\"audit-demo\"}\n```" (no change needed — already correct) - examples/ske-kubeapi-audit-log/README.md:180: "```logql\n{service_name=\"ske\", stackit_log_kind=\"kubernetes-audit\", service_instance_id=\"audit-demo </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ✅ No spelling or grammar issues found. </details> </details> <details> <summary>⚠️ 🏗️ Infrastructure Changes</summary> <details> <summary>🤖 STACKIT Model Serving</summary> - Creates STACKIT SKE cluster with audit enabled and connected to Telemetry Router - Creates STACKIT network for SKE nodes with specified IPv4 prefix and nameservers - Creates STACKIT Observability instance for audit log storage and Grafana access - Creates STACKIT Telemetry Router with access token and destination for Kubernetes audit logs - Creates STACKIT Telemetry Link to route project logs to the Telemetry Router - Deploys Kubernetes namespace and ConfigMap via ephemeral kubeconfig for audit log generation - Exposes outputs for cluster name, router ID/URI, Grafana URL, and kubeconfig command ⚠️ No destructive changes detected. All resources are new creations. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### 1. Identified STACKIT Services The pull request provisions a complete telemetry pipeline to capture and store SKE audit logs. The following services are being introduced: * **SKE (STACKIT Kubernetes Engine):** A new cluster is being provisioned (`stackit_ske_cluster.this`) with audit logging enabled. It includes a node pool using the **Flatcar** OS. * **Network:** A dedicated STACKIT network (`stackit_network.ske_nodes`) is created to host the SKE nodes. * **Observability:** A managed instance (`stackit_observability_instance.audit`) is provisioned to act as the long-term storage and visualization layer (Loki/Grafana). It includes specific credentials for ingestion. * **Telemetry Router:** A managed router instance (`stackit_telemetryrouter_instance.this`) is provisioned to act as the central ingestion and distribution point. * **Telemetry Link:** A project-level link (`stackit_telemetrylink.this`) is established to connect the project's audit stream to the Telemetry Router. **Data Flow Topology:** ```text [ SKE Cluster ] --(Audit Logs)--> [ Telemetry Link ] | v [ Observability ] <--(OTLP)--- [ Telemetry Router ] (Loki/Grafana) ``` ### 2. STACKIT Best Practices & Architectural Review While the implementation is functional, there are several architectural considerations regarding STACKIT best practices: * **Centralized vs. Localized Routing:** The code currently deploys the Telemetry Router within the same project as the SKE cluster. As noted in the documentation, for production environments, it is highly recommended to run a single, central **Telemetry Router** in a dedicated "hub" project and feed it via **Telemetry Links** from various workload projects [“We recommend a single, central Telemetry Router in a core project...”](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/faq/). * **Observability Plan Selection:** The code uses the `Observability-Large-EU01` plan. Ensure this aligns with your expected log volume. Note that the **Observability Large** plan has a specific limit of **1,000 GB** for log storage [“Detailed limits for Observability plans”](https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/). * **Security (ACLs):** The `telemetry_acl` is currently set to `["0.0.0.0/0"]`. This is highly permissive. For a production environment, you should restrict access to the Observability instance to known IP ranges or specific internal networks. * **Observability Ingestion Path:** The `stackit_telemetryrouter_destination` uses the `otlp_http_logs_url` from the Observability instance. This is correct, but ensure that the destination URI in any manual configurations includes the full OTLP-compliant path ending in `/v1/logs`, as the router does not append this automatically [“The destination URI must include the full OTLP-compliant path for logs, ending in /v1/logs”](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/faq/). ### 3. Service Quotas, Constraints, and Limitations Before applying this change, please be aware of the following technical constraints: | Constraint Category | Detail | Impact/Requirement | | :--- | :--- | :--- | | **Telemetry Router Scope** | Only handles **Audit Logs** [“Currently, the service exclusively handles Audit Logs.”](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/basics/overview/) | You cannot use this specific router instance to send application logs or custom metrics. | | **Telemetry Link Uniqueness** | Only one Link per hierarchy object [“Each hierarchy object can only be connected to one Telemetry Router at a time.”](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/faq/) | If this project is already linked to another router, this deployment will fail. | | **Regional Constraint** | Links must target a Router in the same region [“A Telemetry Link can only target a Router in the same region.”](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/faq/) | Ensure `var.stackit_region` is consistent across the SKE cluster, the Router, and the Observability instance. | | **Observability Retention** | Default log retention is **5 days** [“Default retention: 5 days (up to 30 days)”](https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/) | The variable `observability_logs_retention_days` is set to `7`. This is within the allowed range (up to 30 days), but ensure it meets your compliance needs. | | **Observability Quotas** | Storage and sample limits apply [“The number of metrics samples, alerting rules and data storage capacity increase with larger plan sizes.”](https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/) | Monitor your log storage usage. It is recommended to set up alerts when consumption reaches **90%** of your plan's limit [“It is highly recommended to set up early warning alerts... when your resource consumption reaches 90%”](https://docs.stackit.cloud/de/products/logging-and-monitoring/observability/how-tos/how-to-set-up-quota-consumption-alerts/). | </details> </details> <details> <summary>⚠️ 🔒 Security Review</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ No security issues found. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### Security Review Findings Based on the provided git diff, I have identified several security considerations regarding the deployment of the SKE cluster, the Observability instance, and the Telemetry Router. #### 1. STACKIT IAM & Authorization Misconfigurations * 🔴 **High: Overly Permissive Observability Access Control (ACL)** The variable `telemetry_acl` in `020-variables.tf` is defaulted to `["0.0.0.0/0"]`. This is applied to the `stackit_observability_instance` in `040-observability.tf`. In a production environment, allowing open access to your observability/logging endpoint is a significant risk. It is recommended to restrict this to known corporate IP ranges or specific management networks [Enhance The Security Of Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/enhance-the-security-of-your-cluster/) [Enhance The Security Of Your Cluster](https://docs.stackit.cloud/de/products/runtime/kubernetes-engine/how-tos/enhance-the-security-of-your-cluster/). * 🟡 **Medium: Telemetry Router Deployment Context** The code notes that for demo purposes, the router lives in the same project as the cluster [050-telemetry-router.tf]. While acceptable for a proof-of-concept, a production architecture should ideally run the Telemetry Router in a dedicated "hub" project to isolate the telemetry ingestion plane from the workload plane. #### 2. Missing STACKIT-Specific Security Controls * 🔴 **High: Missing SKE API Server Access Control (ACL)** The `stackit_ske_cluster` resource in `070-ske-cluster.tf` does not implement the `extensions.acl` block. Without this, the Kubernetes API server is not restricted by source IP ranges. For production, you should implement the `acl` extension to limit access to trusted CIDRs, such as corporate VPNs or CI/CD egress IPs [Configure Access Control For Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/configure-access-control-for-your-cluster/). * 🟡 **Medium: Lack of Network Policy Baseline** The pull request defines a `stackit_network` for SKE nodes, but there are no Kubernetes `NetworkPolicy` resources defined to control pod-to-pod traffic. By default, pods in a cluster can communicate broadly [Containers & Kubernetes Security — Network Security & Cluster Hardening](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/security-engineer/06-container-&-kubernetes-security/module4.md). I recommend implementing a "default-deny" ingress policy as a starting point for all namespaces [Enhance The Security Of Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/enhance-the-security-of-your-cluster/) [Containers & Kubernetes Security — Network Security & Cluster Hardening](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/security-engineer/06-container-&-kubernetes-security/module4.md). * 🟡 **Medium: Absence of Private Cluster Configuration** The current configuration does not explicitly enable private cluster features. Depending on your security requirements, you may want to evaluate if the API server and node communication should be restricted to private networks to reduce the public attack surface [Enhance The Security Of Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/enhance-the-security-of-your-cluster/). #### 3. Secrets and Sensitive Values * 🟢 **Low: Service Account Key Path** The provider configuration in `010-provider.tf` uses `var.stackit_service_account_key_path`. As long as this variable points to a local file path on the runner and is not hard-coded as a string within the `.tf` files, this follows standard practice. However, ensure the JSON key itself is never committed to version control. #### 4. Compliance and Audit-Logging Considerations * 🟢 **Low: SKE Audit Logging Enabled** The configuration correctly enables SKE audit logging via `audit = { enabled = true }` in `070-ske-cluster.tf`. This is a critical requirement for capturing security-relevant events such as Secret access or RBAC changes [Containers & Kubernetes Security — Secure Operations on Kubernetes](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/security-engineer/06-container-&-kubernetes-security/module5.md). * 🟢 **Low: Centralized Log Routing** The implementation of the `stackit_telemetryrouter_instance` and `stackit_telemetrylink` provides a robust mechanism for centralizing SKE kube-apiserver audit logs into the STACKIT Observability instance. This supports the requirement for centralized logging and long-term retention [Containers & Kubernetes Security — Secure Operations on Kubernetes](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/security-engineer/06-container-&-kubernetes-security/module5.md). --- ### Summary Table | Finding | Severity | Category | Recommendation | | :--- | :--- | :--- | :--- | | Open Observability ACL (`0.0.0.0/0`) | 🔴 High | Authorization | Restrict `telemetry_acl` to trusted CIDRs. | | Missing SKE API ACL | 🔴 High | Network Isolation | Implement `extensions.acl` in `stackit_ske_cluster`. | | Missing Network Policies | 🟡 Medium | Network Isolation | Deploy a default-deny `NetworkPolicy` baseline. | | Single-Project Router | 🟡 Medium | Architecture | Move Telemetry Router to a dedicated hub project for production. | | SKE Audit Logging | ✅ Pass | Compliance | Audit logging is correctly enabled. | </details> </details> <details> <summary>✅ 📐 Example Consistency</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ Example follows repository conventions. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> The provided git diff has been reviewed against the established repository conventions and STACKIT-specific provider requirements. While the example is largely well-structured, there are several deviations regarding repository conventions and provider configuration. ### 🔍 Repository Convention Deviations | Convention | Status | Observation | | :--- | :--- | :--- | | **3-digit numeric prefixes** | ✅ Pass | Files follow the `010-`, `020-` pattern. | | **README.md & MAINTAINERS.md** | ✅ Pass | Both files are present. | | **Variable descriptions & snake_case** | ❌ Fail | `stackit_project_id` in `020-variables.tf` lacks a `description` attribute. | | **Explicit provider version constraints** | ❌ Fail | The `required_providers` block in `010-provider.tf` uses `>=` constraints instead of pinning to an exact version. | | **.terraform.lock.hcl committed** | ✅ Pass | The lock file is present in the diff. | | **Apache 2.0 license headers** | ✅ Pass | Headers are present on all `.tf` files. | #### Fix: Missing Variable Description The variable `stackit_project_id` must include a description to comply with the requirement that all variables have a description attribute. ```hcl variable "stackit_project_id" { type = string description = "The STACKIT project ID." default = "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx" } ``` #### Fix: Provider Version Pinning To ensure environment stability and comply with the requirement for explicit version constraints, avoid using range operators like `>=` in the `required_providers` block. ```hcl terraform { required_version = ">= 1.11.0" required_providers { stackit = { source = "stackitcloud/stackit" version = "0.113.0" } kubernetes = { source = "hashicorp/kubernetes" version = "3.2.1" } } } ``` --- ### ☁️ STACKIT Provider & Usage Observations #### 1. Authentication Method The example uses the **Key flow** via `service_account_key_path` [STACKIT Terraform Provider](https://registry.terraform.io/providers/stackitcloud/stackit/latest/docs). This is the recommended method for most projects [Terraform Fundamentals — Setting Up the Environment and Authentication](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/shared-courses/terraform-fundamentals/module2.md). However, for CI/CD workloads, I recommend transitioning to **Workload Identity Federation (WIF)** to eliminate long-lived secrets [STACKIT Terraform Provider](https://registry.terraform.io/providers/stackitcloud/stackit/latest/docs). #### 2. Beta Resources & Experiments The configuration correctly enables beta resources and specific experiments required for this architecture: * `enable_beta_resources = true` is required to access certain features [Stackit Terraform Provider](https://docs.stackit.cloud/de/developer-tools/stackit-iac/stackit-terraform-provider/) [Stackit Terraform Provider](https://docs.stackit.cloud/developer-tools/stackit-iac/stackit-terraform-provider/). * `experiments = ["ske"]` is explicitly set in the `stackit` provider block to support the `ephemeral` resource `stackit_ske_kubeconfig` [STACKIT Terraform Provider](https://registry.terraform.io/providers/stackitcloud/stackit/latest/docs). #### 3. S3 Backend Note While the example does not implement a remote backend, please note that if you move to an S3 backend using STACKIT Object Storage, setting `use_lockfile = true` currently has no effect [STACKIT Terraform Provider](https://registry.terraform.io/providers/stackitcloud/stackit/latest/docs). For state locking, a **PostgreSQL backend** is the recommended alternative [STACKIT Terraform Provider](https://registry.terraform.io/providers/stackitcloud/stackit/latest/docs). #### 4. SKE Audit Logs (Private Preview) As noted in the `README.md`, the `audit = { enabled = true }` configuration for `stackit_ske_cluster` is currently in **private preview** [README]. Ensure your account is explicitly enabled by STACKIT before attempting to apply this configuration. </details> </details> <details> <summary>✅ 📚 Example README</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ Example READMEs are complete. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### Review of `examples/ske-kubeapi-audit-log` I have reviewed the git diff for the new example directory. Below is my architectural assessment regarding the naming conventions and documentation quality. #### 1. Naming Analysis The directory name `ske-kubeapi-audit-log` is **highly effective** and meets our architectural standards for clarity and accuracy. * **Service Identification:** It explicitly identifies the primary STACKIT service involved (**SKE** - STACKIT Kubernetes Engine). * **Use-Case Specificity:** It clearly defines the technical objective (**kubeapi-audit-log**), which distinguishes it from general SKE cluster provisioning or workload deployment examples. * **Consistency:** The name aligns perfectly with the Terraform resources being introduced (e.g., `stackit_ske_cluster` with `audit = { enabled = true }`) and the logical flow of the telemetry pipeline described in the README. #### 2. README Quality Assessment The `README.md` is **comprehensive and high-quality**, providing significant value for a cloud architect or DevOps engineer looking to implement this pattern. * **Service Explanation:** It clearly outlines the interaction between **SKE**, **Telemetry Router**, **Telemetry Link**, and the **Observability instance** (Loki + Grafana). * **Technical Depth:** It goes beyond a simple "how-to" by explaining the **Layered Filtering** approach (Source $\rightarrow$ Router Instance $\rightarrow$ Destination $\rightarrow$ Query time), which is critical for managing log volume and costs in a production environment. * **Usage Section:** A clear usage section is provided, including the necessary `terraform init` and `terraform apply` commands. * **Value-Add Content:** * It includes a detailed table of **Loki labels** (mapping OpenTelemetry attributes to Loki labels), which is essential for writing effective LogQL queries. * It provides concrete **LogQL examples** for different visibility levels (e.g., filtering out "noise" from gardener/system accounts). * It includes a "What gets created" table that maps Terraform files to specific STACKIT resources, aiding in architectural comprehension. #### Summary Table | Criteria | Status | Architect's Notes | | :--- | :--- | :--- | | **Naming** | ✅ Pass | Clear, service-oriented, and accurately describes the audit-log use case. | | **README Quality** | ✅ Pass | Excellent depth; includes usage, filtering logic, and practical LogQL queries. | ✅ Example READMEs are complete. </details> </details> <details> <summary>✅ 💬 Commit Messages</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ Commit messages are descriptive. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> All provided commit messages are sufficiently descriptive. ✅ Commit messages are descriptive. </details> </details> <details> <summary>✅ 🏷️ Tag Quality</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ README tags accurately reflect the content. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### README Tag Review I have reviewed the `README.md` tags for the new example `examples/ske-kubeapi-audit-log/`. While the tags capture the general intent of the repository, they require adjustments to align with STACKIT product naming conventions and to ensure all relevant managed services are represented. #### Analysis of `examples/ske-kubeapi-audit-log/README.md` The current tag line is: `<!-- tags: ske, kubernetes, audit-log, kube-apiserver, observability, otel, telemetry-router, telemetry-link -->` **1. STACKIT Product Naming & Accuracy** * **`ske`**: Correct. This refers to the **STACKIT Kubernetes Engine** [Ship Traces Logs Metrics To Observability Using Opentelemetry](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/ship-traces-logs-metrics-to-observability-using-opentelemetry/). * **`observability`**: Correct. This refers to the **STACKIT Observability** service [Architecture Of Observability](https://docs.stackit.cloud/de/products/logging-and-monitoring/observability/basics/architecture-of-observability/) [Introduction To Observability](https://docs.stackit.cloud/de/products/logging-and-monitoring/observability/basics/introduction-to-observability/). * **`telemetry-router`**: Correct. This refers to the **STACKIT Telemetry Router** [Architecture](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/basics/architecture/) [Faq](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/faq/). * **`telemetry-link`**: Correct. This refers to the **Telemetry Link** component [Architecture](https://docs.stackit.cloud/de/products/logging-and-monitoring/telemetry-router/basics/architecture/) [Architecture](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/basics/architecture/). * **Missing Product Tags**: The example also utilizes the **STACKIT Network** resource (via `stackit_network`) to provide the node network for the SKE cluster. While "network" is a generic term, in a professional-service repository, explicitly tagging the infrastructure component is beneficial. **2. Open-Source Tools & Patterns** * **`otel`**: While technically correct (referring to OpenTelemetry), the standard convention in this repository for the protocol/pattern is `opentelemetry` to match the service capabilities [Architecture](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/basics/architecture/) [Ship Traces Logs Metrics To Observability Using Opentelemetry](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/ship-traces-logs-metrics-to-observability-using-opentelemetry/). * **`kubernetes` / `audit-log` / `kube-apiserver`**: These are descriptive of the workload/pattern rather than the STACKIT product, which is acceptable for discovery, but they do not conflict with the requirements. **3. Formatting** * All existing tags follow the `lowercase-and-hyphen-separated` rule. #### Summary of Findings | Resource | Current Tag(s) | Status | Recommendation | | :--- | :--- | :--- | :--- | | **SKE Cluster** | `ske` | ✅ | Keep as is. | | **Observability** | `observability` | ✅ | Keep as is. | | **Telemetry Router** | `telemetry-router` | ✅ | Keep as is. | | **Telemetry Link** | `telemetry-link` | ✅ | Keep as is. | | **OpenTelemetry** | `otel` | ⚠️ | Change to `opentelemetry` for consistency with documentation [Ship Traces Logs Metrics To Observability Using Opentelemetry](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/ship-traces-logs-metrics-to-observability-using-opentelemetry/). | | **Network** | *None* | ❌ | Add `network` to reflect the `stackit_network` resource. | #### Suggested Correction To ensure full coverage of the managed resources and consistent naming, please update the first line of `examples/ske-kubeapi-audit-log/README.md` to: ```markdown <!-- tags: ske, kubernetes, audit-log, kube-apiserver, observability, opentelemetry, telemetry-router, telemetry-link, network --> ``` </details> </details> --- _Generated automatically — treat as a hint, not a gate._
chore(agents): generate agents.md
Some checks failed
Default CI / Check README Tags (pull_request) Successful in 1m9s
Default CI / Check for Open TODOs (pull_request) Successful in 1m10s
AI PR Review / AI PR Review (pull_request) Successful in 2m14s
Default CI / Check AGENTS.md is up to date (pull_request) Successful in 48s
Default CI / Pre-Commit Hooks (pull_request) Failing after 1m57s
b309f5b750

🤖 AI PR Review

b309f5b7 · STACKIT Model Serving & STACKIT Cloud Advisor

⚠️ 📝 Spelling & Grammar
🤖 STACKIT Model Serving
  • AGENTS.md:152: "This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the STACKIT Telemetry Router into an Observability instance, where they can be analysed with LogQL in Grafana" → "This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the STACKIT Telemetry Router into an Observability instance, where they can be analyzed with LogQL in Grafana"
  • examples/ske-kubeapi-audit-log/README.md:11: "can be analysed with LogQL in Grafana." → "can be analyzed with LogQL in Grafana."
  • examples/ske-kubeapi-audit-log/README.md:20: "This can be overwhelming so this example aims to show how you can view relevant audit logs as well." → "This can be overwhelming, so this example aims to show how you can view relevant audit logs as well."
  • examples/ske-kubeapi-audit-log/README.md:25: "Please get in contact with STACKIT if you want to enable it." → "Please get in contact with STACKIT if you want to enable it."
  • examples/ske-kubeapi-audit-log/README.md:34: "The SKE audit logs feature is enabled for your account, because this feature is in private review" → "The SKE audit logs feature is enabled for your account, because this feature is in private preview"
  • examples/ske-kubeapi-audit-log/README.md:68: "STACKIT Logs focuses only on providing a managed Loki and Observability includes a managed Grafana on top." → "STACKIT Logs focuses only on providing a managed Loki, and Observability includes a managed Grafana on top."
  • examples/ske-kubeapi-audit-log/README.md:71: "In this case it's easier for testing purposes." → "In this case, it’s easier for testing purposes."
  • examples/ske-kubeapi-audit-log/README.md:86: "The audit policy is managed by SKE and cannot be tuned in any way, so you cannot tell the kube-apiserver to skip watch calls from system:serviceaccount:kube-system:* for example." → "The audit policy is managed by SKE and cannot be tuned in any way, so you cannot tell the kube-apiserver to skip watch calls from system:serviceaccount:kube-system:*, for example."
  • examples/ske-kubeapi-audit-log/README.md:102: "STACKIT sets on every forwarded record (dots become underscores when they arrive as Loki labels):" → "STACKIT sets on every forwarded record (dots become underscores when they arrive as Loki labels):"
  • examples/ske-kubeapi-audit-log/README.md:135: "So on this level we cannot filter on the JSON-formatted content / log line of the kube-apiserver audit logs itself. In this case we just filter on our cluster name and kubernetes-audit." → "So at this level, we cannot filter on the JSON-formatted content / log line of the kube-apiserver audit logs itself. In this case, we just filter on our cluster name and kubernetes-audit."
  • examples/ske-kubeapi-audit-log/README.md:152: "Open the grafana_url output. Get credentials via portal." → "Open the grafana_url output. Get credentials via the portal."
  • examples/ske-kubeapi-audit-log/README.md:163: "logql\n{service_name=\"ske\", stackit_log_kind=\"kubernetes-audit\", service_instance_id=\"audit-demo\"}\n | json user=\"user.username\", verb=\"verb\", res=\"objectRef.resource\", ns=\"objectRef.namespace\", stage=\"stage\"\n | user !~ `system:.*`\n | user !~ `gardener\\.cloud:.*`\n" → "logql\n{service_name=\"ske\", stackit_log_kind=\"kubernetes-audit\", service_instance_id=\"audit-demo\"}\n | json user=\"user.username\", verb=\"verb\", res=\"objectRef.resource\", ns=\"objectRef.namespace\", stage=\"stage\"\n | user !~ `system:.*`\n | user !~ `gardener\\.cloud:.*`\n"
🔍 STACKIT Cloud Advisor

AGENTS.md: line 158: "analysed" -> "analyzed" (Note: While "analysed" is correct in British English, the surrounding documentation and STACKIT technical prose typically follow US English conventions; however, if strictly checking for errors, this is a regional spelling variation rather than a definitive error. Given the context of technical documentation, consistency is key.)

README.md: line 11: "analysed" -> "analyzed"

README.md: line 23: "private review" -> "private preview"

README.md: line 33: "SNA" -> "SNA" (Note: The term "SNA" is used in the text, but the source documentation provided does not define this acronym. As an architect, I recommend ensuring "SNA" is defined earlier in the document or replaced with the full term to avoid ambiguity for new users.)

README.md: line 38: "private review" -> "private preview"

⚠️ 🏗️ Infrastructure Changes
🤖 STACKIT Model Serving
  • Creates new STACKIT provider configuration with SKE experiments enabled
  • Defines variables for SKE cluster, network, and observability configuration
  • Creates STACKIT network resource for SKE nodes with IPv4 prefix and nameservers
  • Creates STACKIT observability instance with Loki+Grafana and associated credential
  • Creates STACKIT telemetry router instance with access token and destination for audit logs
  • Creates STACKIT telemetry link to connect project to telemetry router
  • Creates SKE cluster with audit enabled, using flatcar OS and specific node pool configuration
  • Creates ephemeral kubeconfig for cluster access during Terraform apply
  • Deploys Kubernetes namespace and config map as canary workload to generate audit logs
  • Exposes outputs for cluster name, telemetry router details, Grafana URL, and kubeconfig command
🔍 STACKIT Cloud Advisor

1. Identified STACKIT Services

The pull request provisions a complete telemetry pipeline to capture and store SKE audit logs. The following services are being introduced:

  • SKE (STACKIT Kubernetes Engine): A new cluster is being provisioned with audit.enabled = true to generate the required log stream.
  • Network: A dedicated stackit_network is created to host the SKE nodes.
  • Observability: An stackit_observability_instance is provisioned (using the Observability-Large-EU01 plan) to act as the long-term storage and visualization layer (Loki/Grafana).
  • Telemetry Router: A stackit_telemetryrouter_instance is provisioned to act as the central ingestion and distribution point.
  • Telemetry Link: A stackit_telemetrylink is established to connect the project's audit stream to the Telemetry Router.

Data Flow Architecture:

[ SKE Cluster ] --(Audit Logs)--> [ Telemetry Link ]
                                         |
                                         v
[ Observability ] <--(OTLP)--- [ Telemetry Router ]
(Loki/Grafana)

2. Best Practices & Architectural Review

While the implementation is functional for a demo, several points should be addressed for a production-grade environment:

  • Telemetry Router Topology: The code currently deploys the Telemetry Router within the same project as the SKE cluster. As noted in the code comments and documentation, for production environments, it is a best practice to run a single, central Telemetry Router in a dedicated "hub" project and feed it via Telemetry Links from various workload projects Faq.
  • Observability Plan Selection: The configuration uses the Observability-Large-EU01 plan. This is a high-capacity plan. Ensure the scale of audit logs justifies this, as routing data to internal services incurs fees Overview.
  • Security (ACLs): The telemetry_acl is currently set to ["0.0.0.0/0"]. This is highly permissive. In a production scenario, the Observability instance's ACL should be restricted to known sources or managed via tighter network controls.
  • Log Retention Alignment: The observability_logs_retention_days is set to 7. Note that for Observability plans, the default retention for logs is 5 days, and it can be configured up to 30 days Service Plans Observability. Ensure your compliance requirements align with this setting.

3. Service Quotas, Constraints, and Limitations

Please be aware of the following technical constraints relevant to this deployment:

  • Telemetry Router Scope:
    • Data Types: The Telemetry Router currently exclusively handles Audit Logs. It cannot be used to route custom application logs, metrics, or traces Overview Overview Faq.
    • Regional Constraint: A Telemetry Link can only target a Telemetry Router located in the same region Faq. Since the SKE cluster and Router are both in var.stackit_region, this is currently satisfied.
    • Hierarchy Limitation: A Telemetry Router must always be created within a project; it cannot be created at the folder or organization level Faq.
  • Observability Plan Limits (Large Plan):
  • Observability Data Integrity: There is no backup of telemetry data (metrics, logs, or traces) provided by STACKIT. While configuration files and dashboards are backed up automatically, the actual log data resides in the storage layer and must be managed according to your data protection strategy STACKIT Observability.
  • OTLP Destination Requirements: The configuration uses an OTLP destination. Ensure the destination URI includes the full path ending in /v1/logs (e.g., .../v1/logs), as the Telemetry Router does not automatically append this path Faq.
⚠️ 🔒 Security Review
🤖 STACKIT Model Serving

✅ No security issues found.

🔍 STACKIT Cloud Advisor

1. STACKIT IAM or Authorization Misconfigurations

  • 🔴 High: Overly Permissive Observability Access Control (ACL)
    The variable telemetry_acl in 020-variables.tf is set to ["0.0.0.0/0"]. This allows full access to the Observability instance (including Grafana, Metrics, Logs, and Traces) from any IP address on the internet How To Control Instance Access. For a production environment, this should be restricted to known corporate IP ranges or specific management networks Enhance The Security Of Your Cluster.
  • 🟡 Medium: Potential for Excessive Permissions via Telemetry Link
    The stackit_telemetrylink resource connects the entire project to the telemetry router [060-telemetry-link.tf]. While this is a standard way to stream logs, ensure that the service account used to deploy this infrastructure follows the principle of least privilege, as it manages cross-resource links and telemetry routing.

2. Missing STACKIT-Specific Security Controls

  • 🔴 High: Lack of SKE API Server Access Control (ACL)
    The SKE cluster configuration in 070-ske-cluster.tf does not define an acl block. By default, the Kubernetes API server is exposed to the internet. It is a best practice to restrict API server access to specific, trusted CIDR ranges (e.g., corporate VPN or CI/CD ranges) to reduce the attack surface Enhance The Security Of Your Cluster.
  • 🟡 Medium: Missing Network Policies for SKE Workloads
    While the code creates a kubernetes_namespace_v1 and a kubernetes_config_map_v1 for a canary workload, there are no kubernetes_network_policy resources defined. Without a "default-deny" policy, pods in the cluster may be able to communicate broadly, increasing the risk of lateral movement if a workload is compromised Containers & Kubernetes Security — Network Security & Cluster Hardening.
  • 🟢 Low: Observability Instance Plan Selection
    The observability_plan_name is set to Observability-Large-EU01. While this is a valid plan, ensure the plan size aligns with the expected log volume to prevent performance issues or unexpected costs, though this is more of an operational concern than a direct security gap.

3. Secrets, Credentials, or Sensitive Values

  • 🟡 Medium: Service Account Key Path via Variable
    The provider configuration in 010-provider.tf uses var.stackit_service_account_key_path. While the path itself is not hard-coded, the security of the entire deployment depends on the protection of the JSON key file located at that path. Ensure this file is never committed to version control and is managed via a secure secret management system Containers & Kubernetes Security — Identity, Access & Workload Protection.
  • 🟢 Low: Placeholder Project ID
    The stackit_project_id in 020-variables.tf contains a placeholder (xxxxxxxx-xxxx...). This is acceptable for a template, but ensure the actual value is passed via a sensitive variable or environment variable during runtime.

4. STACKIT Compliance and Audit-Logging Considerations

  • 🟢 Low: Audit Logging is Enabled
    The SKE cluster correctly has audit = { enabled = true } configured in 070-ske-cluster.tf. This is a critical requirement for capturing Kubernetes API activity, such as secret access or RBAC changes, which is essential for incident investigation and compliance Containers & Kubernetes Security — Secure Operations on Kubernetes.
  • 🟢 Low: Log Retention Policy
    The observability_logs_retention_days is set to 7. While this satisfies the technical setup, organizations should ensure this retention period aligns with their specific legal, regulatory, or internal security requirements (e.g., many organizations require 90 days or more) Security Operations & Threat Detection — Security Visibility & Logging.

Summary Architecture of the Proposed Setup

[ Internet ] 
      |
      | (Unrestricted Access via 0.0.0.0/0)
      v
[ STACKIT Observability ] <--- [ Telemetry Router ] <--- [ SKE Cluster (Audit Logs) ]
                                     ^
                                     | (Telemetry Link)
                                     |
                              [ STACKIT Project ]
Security Component Current Status Recommendation
SKE API Access 🔴 Open to Internet Implement Cluster-level ACL Enhance The Security Of Your Cluster
Observability Access 🔴 Open to Internet Restrict telemetry_acl to trusted IPs How To Control Instance Access
Pod Isolation 🟡 No Network Policies Implement Default-Deny NetworkPolicies Containers & Kubernetes Security — Network Security & Cluster Hardening
Audit Logging 🟢 Enabled Maintain and monitor for anomalies Containers & Kubernetes Security — Secure Operations on Kubernetes
⚠️ 📐 Example Consistency
🤖 STACKIT Model Serving

✅ Example follows repository conventions.

🔍 STACKIT Cloud Advisor

The provided git diff has been reviewed against the repository conventions and STACKIT-specific provider requirements. While the example follows most structural conventions, there are critical deviations regarding provider versioning and SKE resource configuration.

🚩 Repository Convention Deviations

1. Missing Variable Descriptions
The variable stackit_project_id in 020-variables.tf lacks a description attribute. All variables must have a description.

variable "stackit_project_id" {
  type        = string
  description = "The STACKIT project ID."
  default     = "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
}

2. Inconsistent Provider Version Constraints
The repository convention requires all providers in required_providers blocks to have an explicit version constraint. In 010-provider.tf, the kubernetes provider uses a minimum version constraint (>= 2.30.0), which is too broad for a strict convention. It should be pinned or constrained more tightly.

terraform {
  required_providers {
    stackit = {
      source  = "stackitcloud/stackit"
      version = ">= 0.113.0"
    }
    kubernetes = {
      source  = "hashicorp/kubernetes"
      version = "~> 3.2.1" 
    }
  }
}

☁️ STACKIT & SKE Specific Concerns

1. SKE Node Pool Configuration (Critical)
In 070-ske-cluster.tf, the node_pools configuration defines max_surge = 1.
According to the SKE FAQ, if a node pool is configured across multiple Availability Zones (e.g., eu01-1, eu01-2, eu01-3), a max_surge value less than the number of configured AZs can result in downtime during rolling updates because the surge capacity is distributed across sub-pools Faq.

To ensure high availability during updates for a multi-AZ pool, max_surge should be increased to match the number of AZs.

# Fix for 070-ske-cluster.tf
node_pools = [
  {
    name               = "standard"
    machine_type       = "c2i.2"
    minimum            = 1
    maximum            = 2
    max_surge          = 3 # Increased to match AZ count to prevent downtime
    availability_zones = ["${var.stackit_region}-1", "${var.stackit_region}-2", "${var.stackit_region}-3"]
    # ... rest of config
  }
]

2. SKE Kubernetes Versioning
The example uses kubernetes_version_min to drive the cluster version. While this is valid, note that SKE automatically manages updates if maintenance.enable_kubernetes_version_updates is set to true stackit_ske_cluster Resource. The use of lifecycle { ignore_changes = [ kubernetes_version_min ] } in 070-ske-cluster.tf is a correct architectural decision to prevent Terraform from fighting the SKE managed update process.

3. Ephemeral Kubeconfig Usage
The use of the ephemeral "stackit_ske_kubeconfig" resource is the recommended pattern for connecting the kubernetes provider to a newly created SKE cluster, as it avoids the need to manage static, long-lived credentials Using Kubernetes Provider with STACKIT SKE.

Summary of Architecture Flow

[ STACKIT Project ]
       |
       +-- [ SKE Cluster ] ----> (Audit Logs) ----+
       |                                         |
       +-- [ Telemetry Router ] <----------------+
                |
                +-- (Filter: kubernetes-audit)
                |
                +-- [ Observability Instance ] (Loki/Grafana)
✅ 📚 Example README
🤖 STACKIT Model Serving

✅ Example READMEs are complete.

🔍 STACKIT Cloud Advisor

Review of examples/ske-kubeapi-audit-log

I have reviewed the git diff for the new example directory. Below is my assessment based on your requirements.

1. Naming

The directory name ske-kubeapi-audit-log is excellent.

  • It clearly identifies the primary STACKIT service involved (SKE).
  • It accurately describes the specific use case (kube-apiserver audit logging).
  • It follows the established naming convention seen in other examples (e.g., ske-velero-backup).

2. README Quality

The README.md is high quality and meets all architectural documentation standards.

  • Service Explanation: It clearly explains the flow of data: SKE \rightarrow Telemetry Router \rightarrow Observability instance (Loki + Grafana). It also provides critical context regarding the "Private Preview" status of the audit = { enabled = true } feature, which is vital for customer expectations.
  • Demonstration Value: The README goes beyond a simple "how-to" by explaining the different layers of filtering (Source, Router Instance, Router Destination, and Query time). This provides significant educational value for an architect designing a logging pipeline.
  • Usage Section: A clear usage section is provided, including:
    • Preparation of terraform.tfvars.
    • Standard workflow: terraform init and terraform apply.
    • Cleanup instructions: terraform destroy.
  • Additional Value: The inclusion of a "Canary Workload" explanation and specific LogQL queries for "Human activity only" makes the example immediately testable and verifiable for the user.

Summary Table

Criteria Status Notes
Naming ✅ Pass Clear, accurate, and follows convention.
README Content ✅ Pass Explains services, use-case, and filtering logic.
README Usage ✅ Pass Includes init, apply, and destroy.

✅ Example READMEs are complete.

✅ 💬 Commit Messages
🤖 STACKIT Model Serving

✅ Commit messages are descriptive.

🔍 STACKIT Cloud Advisor

Based on my review of the provided commit messages, I have identified one entry that is too vague to provide meaningful context for a professional services repository. While the other messages clearly state the scope (e.g., ske for Stackit Kubernetes Engine) and the specific action taken, the first message lacks sufficient detail.

Commit Message Quality Audit

Current Message Status Assessment Suggested Alternative
chore(agents): generate agents.md ⚠️ Vague While it follows a conventional commit format, "generate agents.md" does not explain why the file was generated or what specific agent configurations/documentation it contains. docs(agents): generate documentation for agent deployment configurations
example(ske): fix telemetry router destination filter level and keys ✅ Descriptive Clearly identifies the service (SKE), the component (telemetry router), and the specific technical fix (filter level and keys). N/A
example(ske): add example for kubeapi audit logs ✅ Descriptive Clearly identifies the service (SKE) and the specific resource/feature being added (kubeapi audit logs). N/A

Architect's Recommendation for Improvement

To maintain a high-quality repository, especially within professional services where these commits often serve as a reference for customers, I recommend adhering to the following principles:

  1. Contextualize the "Why": A commit should not just say what changed (e.g., "generate file"), but briefly imply why (e.g., "to document new agent capabilities").
  2. Service-Centric Scoping: Continue using the scope parenthesis (e.g., ske, workflows, notebooks) to allow developers to quickly filter history by STACKIT service.
  3. Granularity: As noted in the STACKIT Git best practices, keeping commits small and focused makes them easier to revert or adapt Set Up A Git Repository. Ensure that the descriptive nature of the message matches the granularity of the change.
⚠️ 🏷️ Tag Quality
🤖 STACKIT Model Serving

✅ README tags accurately reflect the content.

🔍 STACKIT Cloud Advisor

README Tag Review: examples/ske-kubeapi-audit-log/

I have reviewed the tags in the newly introduced example examples/ske-kubeapi-audit-log/. While the tags capture the general intent, there are several discrepancies regarding STACKIT product naming conventions and missing technical patterns that should be addressed to ensure consistency with our professional service standards.

1. STACKIT Product Tagging

The current tags include otel and telemetry-router, which are technical components, but they miss the official STACKIT product names required for accurate resource categorization.

  • Issue: The tags use observability and telemetry-router, but do not explicitly use the STACKIT product identifiers for the core services being orchestrated.
  • Requirement: Ensure the official product names (e.g., ske, observability) are present.
  • Correction: The tag line should be updated to include the specific STACKIT service names.

2. Missing Open-Source Tools & Patterns

The example demonstrates a sophisticated telemetry pipeline using OpenTelemetry (OTel) standards to bridge SKE and Observability.

  • Issue: While otel is present, the specific pattern of using OpenTelemetry as a standard for shipping logs is a key architectural detail that should be explicitly tagged to help users find pattern-based examples.
  • Observation: The example also utilizes kubernetes (via the provider), which is correctly tagged.

3. Tag Formatting & Consistency

  • Status: All existing tags are lowercase and hyphen-separated. This complies with our internal style guide.

Summary of Findings & Recommendations

Resource Current Tags Status Suggested Correction
examples/ske-kubeapi-audit-log/README.md <!-- tags: ske, kubernetes, audit-log, kube-apiserver, observability, otel, telemetry-router, telemetry-link --> ⚠️ Incomplete <!-- tags: ske, observability, telemetry-router, telemetry-link, kubernetes, opentelemetry, audit-log -->

Architectural Context

To visualize why these tags are critical, here is the data flow being documented in this example:

[ SKE Cluster ] --(Audit Logs)--> [ Telemetry Link ]
                                        |
                                        v
[ Observability ] <--(OTLP/Loki)-- [ Telemetry Router ]
      (Grafana)

Architect's Note:
I recommend updating the tags to include opentelemetry instead of just otel to align with the full name of the standard used in the stackit_telemetryrouter_destination configuration, and ensuring observability is treated as the primary product tag.


Generated automatically — treat as a hint, not a gate.

## 🤖 AI PR Review > [`b309f5b7`](https://professional-service.git.onstackit.cloud/professional-service-best-practices/professional-service/commit/b309f5b7505cc6ecd61eedff2cfcb7dbd4bdc021) · STACKIT Model Serving & STACKIT Cloud Advisor <details> <summary>⚠️ 📝 Spelling & Grammar</summary> <details> <summary>🤖 STACKIT Model Serving</summary> - AGENTS.md:152: "This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the **STACKIT Telemetry Router** into an **Observability instance**, where they can be analysed with LogQL in Grafana" → "This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the **STACKIT Telemetry Router** into an **Observability instance**, where they can be analyzed with LogQL in Grafana" - examples/ske-kubeapi-audit-log/README.md:11: "can be analysed with LogQL in Grafana." → "can be analyzed with LogQL in Grafana." - examples/ske-kubeapi-audit-log/README.md:20: "This can be overwhelming so this example aims to show how you can view relevant audit logs as well." → "This can be overwhelming, so this example aims to show how you can view relevant audit logs as well." - examples/ske-kubeapi-audit-log/README.md:25: "Please get in contact with STACKIT if you want to enable it." → "Please get in contact with STACKIT if you want to enable it." - examples/ske-kubeapi-audit-log/README.md:34: "The SKE audit logs feature is enabled for your account, because this feature is in private review" → "The SKE audit logs feature is enabled for your account, because this feature is in private preview" - examples/ske-kubeapi-audit-log/README.md:68: "STACKIT Logs focuses only on providing a managed Loki and Observability includes a managed Grafana on top." → "STACKIT Logs focuses only on providing a managed Loki, and Observability includes a managed Grafana on top." - examples/ske-kubeapi-audit-log/README.md:71: "In this case it's easier for testing purposes." → "In this case, it’s easier for testing purposes." - examples/ske-kubeapi-audit-log/README.md:86: "The audit policy is managed by SKE and cannot be tuned in any way, so you cannot tell the kube-apiserver to skip `watch` calls from `system:serviceaccount:kube-system:*` for example." → "The audit policy is managed by SKE and cannot be tuned in any way, so you cannot tell the kube-apiserver to skip `watch` calls from `system:serviceaccount:kube-system:*`, for example." - examples/ske-kubeapi-audit-log/README.md:102: "STACKIT sets on every forwarded record (dots become underscores when they arrive as Loki labels):" → "STACKIT sets on every forwarded record (dots become underscores when they arrive as Loki labels):" - examples/ske-kubeapi-audit-log/README.md:135: "So on this level we cannot filter on the JSON-formatted content / log line of the kube-apiserver audit logs itself. In this case we just filter on our cluster name and `kubernetes-audit`." → "So at this level, we cannot filter on the JSON-formatted content / log line of the kube-apiserver audit logs itself. In this case, we just filter on our cluster name and `kubernetes-audit`." - examples/ske-kubeapi-audit-log/README.md:152: "Open the `grafana_url` output. Get credentials via portal." → "Open the `grafana_url` output. Get credentials via the portal." - examples/ske-kubeapi-audit-log/README.md:163: "```logql\n{service_name=\"ske\", stackit_log_kind=\"kubernetes-audit\", service_instance_id=\"audit-demo\"}\n | json user=\"user.username\", verb=\"verb\", res=\"objectRef.resource\", ns=\"objectRef.namespace\", stage=\"stage\"\n | user !~ `system:.*`\n | user !~ `gardener\\.cloud:.*`\n```" → "```logql\n{service_name=\"ske\", stackit_log_kind=\"kubernetes-audit\", service_instance_id=\"audit-demo\"}\n | json user=\"user.username\", verb=\"verb\", res=\"objectRef.resource\", ns=\"objectRef.namespace\", stage=\"stage\"\n | user !~ `system:.*`\n | user !~ `gardener\\.cloud:.*`\n```" </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> AGENTS.md: line 158: "analysed" -> "analyzed" (Note: While "analysed" is correct in British English, the surrounding documentation and STACKIT technical prose typically follow US English conventions; however, if strictly checking for errors, this is a regional spelling variation rather than a definitive error. Given the context of technical documentation, consistency is key.) README.md: line 11: "analysed" -> "analyzed" README.md: line 23: "private review" -> "private preview" README.md: line 33: "SNA" -> "SNA" (Note: The term "SNA" is used in the text, but the source documentation provided does not define this acronym. As an architect, I recommend ensuring "SNA" is defined earlier in the document or replaced with the full term to avoid ambiguity for new users.) README.md: line 38: "private review" -> "private preview" </details> </details> <details> <summary>⚠️ 🏗️ Infrastructure Changes</summary> <details> <summary>🤖 STACKIT Model Serving</summary> - Creates new STACKIT provider configuration with SKE experiments enabled - Defines variables for SKE cluster, network, and observability configuration - Creates STACKIT network resource for SKE nodes with IPv4 prefix and nameservers - Creates STACKIT observability instance with Loki+Grafana and associated credential - Creates STACKIT telemetry router instance with access token and destination for audit logs - Creates STACKIT telemetry link to connect project to telemetry router - Creates SKE cluster with audit enabled, using flatcar OS and specific node pool configuration - Creates ephemeral kubeconfig for cluster access during Terraform apply - Deploys Kubernetes namespace and config map as canary workload to generate audit logs - Exposes outputs for cluster name, telemetry router details, Grafana URL, and kubeconfig command </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### 1. Identified STACKIT Services The pull request provisions a complete telemetry pipeline to capture and store SKE audit logs. The following services are being introduced: * **SKE (STACKIT Kubernetes Engine):** A new cluster is being provisioned with `audit.enabled = true` to generate the required log stream. * **Network:** A dedicated `stackit_network` is created to host the SKE nodes. * **Observability:** An `stackit_observability_instance` is provisioned (using the `Observability-Large-EU01` plan) to act as the long-term storage and visualization layer (Loki/Grafana). * **Telemetry Router:** A `stackit_telemetryrouter_instance` is provisioned to act as the central ingestion and distribution point. * **Telemetry Link:** A `stackit_telemetrylink` is established to connect the project's audit stream to the Telemetry Router. **Data Flow Architecture:** ```text [ SKE Cluster ] --(Audit Logs)--> [ Telemetry Link ] | v [ Observability ] <--(OTLP)--- [ Telemetry Router ] (Loki/Grafana) ``` --- ### 2. Best Practices & Architectural Review While the implementation is functional for a demo, several points should be addressed for a production-grade environment: * **Telemetry Router Topology:** The code currently deploys the Telemetry Router within the same project as the SKE cluster. As noted in the code comments and documentation, for production environments, it is a best practice to run a single, central **Telemetry Router** in a dedicated "hub" project and feed it via **Telemetry Links** from various workload projects [Faq](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/faq/). * **Observability Plan Selection:** The configuration uses the `Observability-Large-EU01` plan. This is a high-capacity plan. Ensure the scale of audit logs justifies this, as routing data to internal services incurs fees [Overview](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/basics/overview/). * **Security (ACLs):** The `telemetry_acl` is currently set to `["0.0.0.0/0"]`. This is highly permissive. In a production scenario, the Observability instance's ACL should be restricted to known sources or managed via tighter network controls. * **Log Retention Alignment:** The `observability_logs_retention_days` is set to `7`. Note that for **Observability** plans, the default retention for logs is **5 days**, and it can be configured up to **30 days** [Service Plans Observability](https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/). Ensure your compliance requirements align with this setting. --- ### 3. Service Quotas, Constraints, and Limitations Please be aware of the following technical constraints relevant to this deployment: * **Telemetry Router Scope:** * **Data Types:** The Telemetry Router currently exclusively handles **Audit Logs**. It cannot be used to route custom application logs, metrics, or traces [Overview](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/basics/overview/) [Overview](https://docs.stackit.cloud/de/products/logging-and-monitoring/telemetry-router/basics/overview/) [Faq](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/faq/). * **Regional Constraint:** A **Telemetry Link** can only target a **Telemetry Router** located in the same region [Faq](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/faq/). Since the SKE cluster and Router are both in `var.stackit_region`, this is currently satisfied. * **Hierarchy Limitation:** A **Telemetry Router** must always be created within a **project**; it cannot be created at the folder or organization level [Faq](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/faq/). * **Observability Plan Limits (Large Plan):** * By selecting the **Large** plan, you are subject to specific quotas: * **Logs/Traces Size:** Up to **1,000 GB** [Service Plans Observability](https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/). * **Logs/Traces Retention:** Default **5 days**, configurable up to **30 days** [Service Plans Observability](https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/). * **Metrics Samples:** Up to **300,000** per minute [Service Plans Observability](https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/). * **Observability Data Integrity:** There is **no backup** of telemetry data (metrics, logs, or traces) provided by STACKIT. While configuration files and dashboards are backed up automatically, the actual log data resides in the storage layer and must be managed according to your data protection strategy [STACKIT Observability](https://stackit.com/en/asset/download/34321/file/Service_certificate_STACKIT_Observability.pdf?version=10). * **OTLP Destination Requirements:** The configuration uses an OTLP destination. Ensure the destination URI includes the full path ending in `/v1/logs` (e.g., `.../v1/logs`), as the Telemetry Router does not automatically append this path [Faq](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/faq/). </details> </details> <details> <summary>⚠️ 🔒 Security Review</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ No security issues found. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### 1. STACKIT IAM or Authorization Misconfigurations * 🔴 **High: Overly Permissive Observability Access Control (ACL)** The variable `telemetry_acl` in `020-variables.tf` is set to `["0.0.0.0/0"]`. This allows full access to the Observability instance (including Grafana, Metrics, Logs, and Traces) from any IP address on the internet [How To Control Instance Access](https://docs.stackit.cloud/products/logging-and-monitoring/observability/how-tos/how-to-control-instance-access/). For a production environment, this should be restricted to known corporate IP ranges or specific management networks [Enhance The Security Of Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/enhance-the-security-of-your-cluster/). * 🟡 **Medium: Potential for Excessive Permissions via Telemetry Link** The `stackit_telemetrylink` resource connects the entire project to the telemetry router [060-telemetry-link.tf]. While this is a standard way to stream logs, ensure that the service account used to deploy this infrastructure follows the principle of least privilege, as it manages cross-resource links and telemetry routing. ### 2. Missing STACKIT-Specific Security Controls * 🔴 **High: Lack of SKE API Server Access Control (ACL)** The SKE cluster configuration in `070-ske-cluster.tf` does not define an `acl` block. By default, the Kubernetes API server is exposed to the internet. It is a best practice to restrict API server access to specific, trusted CIDR ranges (e.g., corporate VPN or CI/CD ranges) to reduce the attack surface [Enhance The Security Of Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/enhance-the-security-of-your-cluster/). * 🟡 **Medium: Missing Network Policies for SKE Workloads** While the code creates a `kubernetes_namespace_v1` and a `kubernetes_config_map_v1` for a canary workload, there are no `kubernetes_network_policy` resources defined. Without a "default-deny" policy, pods in the cluster may be able to communicate broadly, increasing the risk of lateral movement if a workload is compromised [Containers & Kubernetes Security — Network Security & Cluster Hardening](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/security-engineer/06-container-&-kubernetes-security/module4.md). * 🟢 **Low: Observability Instance Plan Selection** The `observability_plan_name` is set to `Observability-Large-EU01`. While this is a valid plan, ensure the plan size aligns with the expected log volume to prevent performance issues or unexpected costs, though this is more of an operational concern than a direct security gap. ### 3. Secrets, Credentials, or Sensitive Values * 🟡 **Medium: Service Account Key Path via Variable** The provider configuration in `010-provider.tf` uses `var.stackit_service_account_key_path`. While the path itself is not hard-coded, the security of the entire deployment depends on the protection of the JSON key file located at that path. Ensure this file is never committed to version control and is managed via a secure secret management system [Containers & Kubernetes Security — Identity, Access & Workload Protection](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/security-engineer/06-container-&-kubernetes-security/module3.md). * 🟢 **Low: Placeholder Project ID** The `stackit_project_id` in `020-variables.tf` contains a placeholder (`xxxxxxxx-xxxx...`). This is acceptable for a template, but ensure the actual value is passed via a sensitive variable or environment variable during runtime. ### 4. STACKIT Compliance and Audit-Logging Considerations * 🟢 **Low: Audit Logging is Enabled** The SKE cluster correctly has `audit = { enabled = true }` configured in `070-ske-cluster.tf`. This is a critical requirement for capturing Kubernetes API activity, such as secret access or RBAC changes, which is essential for incident investigation and compliance [Containers & Kubernetes Security — Secure Operations on Kubernetes](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/security-engineer/06-container-&-kubernetes-security/module5.md). * 🟢 **Low: Log Retention Policy** The `observability_logs_retention_days` is set to `7`. While this satisfies the technical setup, organizations should ensure this retention period aligns with their specific legal, regulatory, or internal security requirements (e.g., many organizations require 90 days or more) [Security Operations & Threat Detection — Security Visibility & Logging](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/security-engineer/05-security-operations-&-threat-detection/module1.md). ### Summary Architecture of the Proposed Setup ```text [ Internet ] | | (Unrestricted Access via 0.0.0.0/0) v [ STACKIT Observability ] <--- [ Telemetry Router ] <--- [ SKE Cluster (Audit Logs) ] ^ | (Telemetry Link) | [ STACKIT Project ] ``` | Security Component | Current Status | Recommendation | | :--- | :--- | :--- | | **SKE API Access** | 🔴 Open to Internet | Implement Cluster-level ACL [Enhance The Security Of Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/enhance-the-security-of-your-cluster/) | | **Observability Access** | 🔴 Open to Internet | Restrict `telemetry_acl` to trusted IPs [How To Control Instance Access](https://docs.stackit.cloud/products/logging-and-monitoring/observability/how-tos/how-to-control-instance-access/) | | **Pod Isolation** | 🟡 No Network Policies | Implement Default-Deny NetworkPolicies [Containers & Kubernetes Security — Network Security & Cluster Hardening](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/security-engineer/06-container-&-kubernetes-security/module4.md) | | **Audit Logging** | 🟢 Enabled | Maintain and monitor for anomalies [Containers & Kubernetes Security — Secure Operations on Kubernetes](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/security-engineer/06-container-&-kubernetes-security/module5.md) | </details> </details> <details> <summary>⚠️ 📐 Example Consistency</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ Example follows repository conventions. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> The provided git diff has been reviewed against the repository conventions and STACKIT-specific provider requirements. While the example follows most structural conventions, there are critical deviations regarding provider versioning and SKE resource configuration. ### 🚩 Repository Convention Deviations **1. Missing Variable Descriptions** The variable `stackit_project_id` in `020-variables.tf` lacks a `description` attribute. All variables must have a description. ```hcl variable "stackit_project_id" { type = string description = "The STACKIT project ID." default = "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx" } ``` **2. Inconsistent Provider Version Constraints** The repository convention requires all providers in `required_providers` blocks to have an **explicit version constraint**. In `010-provider.tf`, the `kubernetes` provider uses a minimum version constraint (`>= 2.30.0`), which is too broad for a strict convention. It should be pinned or constrained more tightly. ```hcl terraform { required_providers { stackit = { source = "stackitcloud/stackit" version = ">= 0.113.0" } kubernetes = { source = "hashicorp/kubernetes" version = "~> 3.2.1" } } } ``` --- ### ☁️ STACKIT & SKE Specific Concerns **1. SKE Node Pool Configuration (Critical)** In `070-ske-cluster.tf`, the `node_pools` configuration defines `max_surge = 1`. According to the **SKE FAQ**, if a node pool is configured across multiple Availability Zones (e.g., `eu01-1`, `eu01-2`, `eu01-3`), a `max_surge` value less than the number of configured AZs can result in downtime during rolling updates because the surge capacity is distributed across sub-pools [Faq](https://docs.stackit.cloud/products/runtime/kubernetes-engine/faq/). To ensure high availability during updates for a multi-AZ pool, `max_surge` should be increased to match the number of AZs. ```hcl # Fix for 070-ske-cluster.tf node_pools = [ { name = "standard" machine_type = "c2i.2" minimum = 1 maximum = 2 max_surge = 3 # Increased to match AZ count to prevent downtime availability_zones = ["${var.stackit_region}-1", "${var.stackit_region}-2", "${var.stackit_region}-3"] # ... rest of config } ] ``` **2. SKE Kubernetes Versioning** The example uses `kubernetes_version_min` to drive the cluster version. While this is valid, note that SKE automatically manages updates if `maintenance.enable_kubernetes_version_updates` is set to `true` [stackit_ske_cluster Resource](https://registry.terraform.io/providers/stackitcloud/stackit/latest/docs/resources/ske_cluster). The use of `lifecycle { ignore_changes = [ kubernetes_version_min ] }` in `070-ske-cluster.tf` is a correct architectural decision to prevent Terraform from fighting the SKE managed update process. **3. Ephemeral Kubeconfig Usage** The use of the `ephemeral "stackit_ske_kubeconfig"` resource is the recommended pattern for connecting the `kubernetes` provider to a newly created SKE cluster, as it avoids the need to manage static, long-lived credentials [Using Kubernetes Provider with STACKIT SKE](https://registry.terraform.io/providers/stackitcloud/stackit/latest/docs/guides/kubernetes_provider_ske). ### Summary of Architecture Flow ```ascii [ STACKIT Project ] | +-- [ SKE Cluster ] ----> (Audit Logs) ----+ | | +-- [ Telemetry Router ] <----------------+ | +-- (Filter: kubernetes-audit) | +-- [ Observability Instance ] (Loki/Grafana) ``` </details> </details> <details> <summary>✅ 📚 Example README</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ Example READMEs are complete. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### Review of `examples/ske-kubeapi-audit-log` I have reviewed the git diff for the new example directory. Below is my assessment based on your requirements. #### 1. Naming The directory name `ske-kubeapi-audit-log` is **excellent**. - It clearly identifies the primary STACKIT service involved (**SKE**). - It accurately describes the specific use case (**kube-apiserver audit logging**). - It follows the established naming convention seen in other examples (e.g., `ske-velero-backup`). #### 2. README Quality The `README.md` is **high quality** and meets all architectural documentation standards. * **Service Explanation:** It clearly explains the flow of data: SKE $\rightarrow$ Telemetry Router $\rightarrow$ Observability instance (Loki + Grafana). It also provides critical context regarding the "Private Preview" status of the `audit = { enabled = true }` feature, which is vital for customer expectations. * **Demonstration Value:** The README goes beyond a simple "how-to" by explaining the different layers of filtering (Source, Router Instance, Router Destination, and Query time). This provides significant educational value for an architect designing a logging pipeline. * **Usage Section:** A clear usage section is provided, including: * Preparation of `terraform.tfvars`. * Standard workflow: `terraform init` and `terraform apply`. * Cleanup instructions: `terraform destroy`. * **Additional Value:** The inclusion of a "Canary Workload" explanation and specific LogQL queries for "Human activity only" makes the example immediately testable and verifiable for the user. #### Summary Table | Criteria | Status | Notes | | :--- | :--- | :--- | | **Naming** | ✅ Pass | Clear, accurate, and follows convention. | | **README Content** | ✅ Pass | Explains services, use-case, and filtering logic. | | **README Usage** | ✅ Pass | Includes `init`, `apply`, and `destroy`. | ✅ Example READMEs are complete. </details> </details> <details> <summary>✅ 💬 Commit Messages</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ Commit messages are descriptive. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> Based on my review of the provided commit messages, I have identified one entry that is too vague to provide meaningful context for a professional services repository. While the other messages clearly state the scope (e.g., `ske` for Stackit Kubernetes Engine) and the specific action taken, the first message lacks sufficient detail. ### Commit Message Quality Audit | Current Message | Status | Assessment | Suggested Alternative | | :--- | :--- | :--- | :--- | | `chore(agents): generate agents.md` | ⚠️ **Vague** | While it follows a conventional commit format, "generate agents.md" does not explain *why* the file was generated or what specific agent configurations/documentation it contains. | `docs(agents): generate documentation for agent deployment configurations` | | `example(ske): fix telemetry router destination filter level and keys` | ✅ **Descriptive** | Clearly identifies the service (**SKE**), the component (telemetry router), and the specific technical fix (filter level and keys). | N/A | | `example(ske): add example for kubeapi audit logs` | ✅ **Descriptive** | Clearly identifies the service (**SKE**) and the specific resource/feature being added (kubeapi audit logs). | N/A | ### Architect's Recommendation for Improvement To maintain a high-quality repository, especially within professional services where these commits often serve as a reference for customers, I recommend adhering to the following principles: 1. **Contextualize the "Why":** A commit should not just say *what* changed (e.g., "generate file"), but briefly imply *why* (e.g., "to document new agent capabilities"). 2. **Service-Centric Scoping:** Continue using the scope parenthesis (e.g., `ske`, `workflows`, `notebooks`) to allow developers to quickly filter history by STACKIT service. 3. **Granularity:** As noted in the STACKIT Git best practices, keeping commits small and focused makes them easier to revert or adapt [Set Up A Git Repository](https://docs.stackit.cloud/products/developer-platform/git/how-tos/set-up-a-git-repository/). Ensure that the descriptive nature of the message matches the granularity of the change. </details> </details> <details> <summary>⚠️ 🏷️ Tag Quality</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ README tags accurately reflect the content. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### README Tag Review: `examples/ske-kubeapi-audit-log/` I have reviewed the tags in the newly introduced example `examples/ske-kubeapi-audit-log/`. While the tags capture the general intent, there are several discrepancies regarding STACKIT product naming conventions and missing technical patterns that should be addressed to ensure consistency with our professional service standards. #### 1. STACKIT Product Tagging The current tags include `otel` and `telemetry-router`, which are technical components, but they miss the official STACKIT product names required for accurate resource categorization. * **Issue:** The tags use `observability` and `telemetry-router`, but do not explicitly use the STACKIT product identifiers for the core services being orchestrated. * **Requirement:** Ensure the official product names (e.g., `ske`, `observability`) are present. * **Correction:** The tag line should be updated to include the specific STACKIT service names. #### 2. Missing Open-Source Tools & Patterns The example demonstrates a sophisticated telemetry pipeline using OpenTelemetry (OTel) standards to bridge SKE and Observability. * **Issue:** While `otel` is present, the specific pattern of using **OpenTelemetry** as a standard for shipping logs is a key architectural detail that should be explicitly tagged to help users find pattern-based examples. * **Observation:** The example also utilizes `kubernetes` (via the provider), which is correctly tagged. #### 3. Tag Formatting & Consistency * **Status:** All existing tags are lowercase and hyphen-separated. This complies with our internal style guide. --- ### Summary of Findings & Recommendations | Resource | Current Tags | Status | Suggested Correction | | :--- | :--- | :--- | :--- | | `examples/ske-kubeapi-audit-log/README.md` | `<!-- tags: ske, kubernetes, audit-log, kube-apiserver, observability, otel, telemetry-router, telemetry-link -->` | ⚠️ **Incomplete** | `<!-- tags: ske, observability, telemetry-router, telemetry-link, kubernetes, opentelemetry, audit-log -->` | #### Architectural Context To visualize why these tags are critical, here is the data flow being documented in this example: ```text [ SKE Cluster ] --(Audit Logs)--> [ Telemetry Link ] | v [ Observability ] <--(OTLP/Loki)-- [ Telemetry Router ] (Grafana) ``` **Architect's Note:** I recommend updating the tags to include `opentelemetry` instead of just `otel` to align with the full name of the standard used in the `stackit_telemetryrouter_destination` configuration, and ensuring `observability` is treated as the primary product tag. </details> </details> --- _Generated automatically — treat as a hint, not a gate._
chore: fix pre-commit
All checks were successful
Default CI / Check README Tags (pull_request) Successful in 50s
Default CI / Check for Open TODOs (pull_request) Successful in 1m3s
Default CI / Pre-Commit Hooks (pull_request) Successful in 2m0s
AI PR Review / AI PR Review (pull_request) Successful in 2m4s
Default CI / Check AGENTS.md is up to date (pull_request) Successful in 1m1s
b8e1ba67da

🤖 AI PR Review

b8e1ba67 · STACKIT Model Serving & STACKIT Cloud Advisor

⚠️ 📝 Spelling & Grammar
🤖 STACKIT Model Serving
  • AGENTS.md:152: “This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the STACKIT Telemetry Router into an Observability instance, where they can be analysed with LogQL in Grafana” → “This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the STACKIT Telemetry Router into an Observability instance, where they can be analyzed with LogQL in Grafana”
  • examples/ske-kubeapi-audit-log/README.md:14: “This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the STACKIT Telemetry Router into an Observability instance, where they can be analysed with LogQL in Grafana.” → “This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the STACKIT Telemetry Router into an Observability instance, where they can be analyzed with LogQL in Grafana.”
  • examples/ske-kubeapi-audit-log/README.md:21: “A lot of logs are produced by the cluster itself, doing get/list/watch calls coming from kublet or gardener (infrastructure to manage your SKE). This can be overwhelming so this example aims to show how you can view relevant audit logs as well.” → “A lot of logs are produced by the cluster itself, doing get/list/watch calls coming from kubelet or gardener (infrastructure to manage your SKE). This can be overwhelming, so this example aims to show how you can view relevant audit logs as well.”
  • examples/ske-kubeapi-audit-log/README.md:32: “Please get in contact with STACKIT if you want to enable it.” → “Please get in contact with STACKIT if you want to enable it.”
  • examples/ske-kubeapi-audit-log/README.md:41: “The SKE audit logs feature is enabled for your account, because this feature is in private review” → “The SKE audit logs feature is enabled for your account, because this feature is in private preview”
  • examples/ske-kubeapi-audit-log/README.md:87: “STACKIT Logs focuses only on providing a managed Loki and Observability includes a managed Grafana on top.” → “STACKIT Logs focuses only on providing a managed Loki, and Observability includes a managed Grafana on top.”
  • examples/ske-kubeapi-audit-log/README.md:131: “So on this level we cannot filter on the JSON-formatted content / log line of the kube-apiserver audit logs itself.” → “So at this level, we cannot filter on the JSON-formatted content / log line of the kube-apiserver audit logs themselves.”
  • examples/ske-kubeapi-audit-log/README.md:144: “Open the grafana_url output. Get credentials via portal.” → “Open the grafana_url output. Get credentials via the portal.”
  • examples/ske-kubeapi-audit-log/README.md:165: “The apply already produces findable audit records: the canary Namespace and ConfigMap are created” → “The apply already produces findable audit records: the canary Namespace and ConfigMap are created.”
🔍 STACKIT Cloud Advisor

✅ No spelling or grammar issues found.

⚠️ 🏗️ Infrastructure Changes
🤖 STACKIT Model Serving
  • Creates new STACKIT network for SKE nodes with IPv4 prefix and nameservers
  • Deploys STACKIT Observability instance (Loki+Grafana) with configured log retention and ACL
  • Creates STACKIT Telemetry Router instance and associated access token
  • Configures Telemetry Router destination to forward Kubernetes audit logs to Observability
  • Establishes Telemetry Link between project and Telemetry Router
  • Deploys SKE Kubernetes cluster with audit logging enabled and connected to Telemetry Router
  • Creates ephemeral kubeconfig for cluster access during Terraform apply
  • Deploys Kubernetes namespace and ConfigMap as audit canary workload
  • Exposes outputs for cluster name, router details, Grafana URL, and kubeconfig command
🔍 STACKIT Cloud Advisor

1. Identified STACKIT Services

The pull request provisions a complete telemetry pipeline to capture and store SKE audit logs. The following services are being introduced:

  • SKE (STACKIT Kubernetes Engine): A new cluster is being provisioned with audit.enabled = true to generate the required log stream.
  • Network: A dedicated stackit_network is created to host the SKE nodes.
  • Observability: An stackit_observability_instance is provisioned (using the Observability-Large-EU01 plan) to act as the long-term storage and visualization layer (Loki/Grafana).
  • Telemetry Router: An stackit_telemetryrouter_instance is provisioned to act as the central ingestion and distribution point.
  • Telemetry Link: A stackit_telemetrylink is created to connect the project's audit stream to the Telemetry Router.

Data Flow Architecture:

[ SKE Cluster ] --(Audit Logs)--> [ Telemetry Link ]
                                         |
                                         v
[ Observability ] <--(OTLP)--- [ Telemetry Router ]
(Loki/Grafana)

2. Best Practices & Architectural Review

While the implementation is functional, there are several architectural considerations regarding STACKIT best practices:

  • Router Topology (Hub-and-Spoke): The code currently deploys the Telemetry Router within the same project as the SKE cluster. As noted in the code comments and documentation, for production environments, it is a best practice to run a single, central Telemetry Router in a dedicated "hub" project and feed it via Telemetry Links from various workload projects Faq.
  • Observability Plan Selection: The configuration uses the Observability-Large-EU01 plan. This is a high-capacity plan providing up to 1,000 GB of log storage and 300,000 metrics samples per minute Service Plans Observability. Ensure this aligns with your actual volume to avoid unnecessary costs.
  • Security & ACLs: The telemetry_acl is currently set to ["0.0.0.0/0"]. This is highly permissive. For a production observability instance, you should restrict this ACL to known source IP ranges or specific network segments.
  • Log Retention: The observability_logs_retention_days is set to 7. Note that for the Observability plan group, the default retention is 5 days, and it can be configured up to 30 days Service Plans Observability.

3. Service Quotas, Constraints, and Limitations

Please be aware of the following technical constraints identified in the service documentation:

  • Telemetry Router Scope:
    • Data Types: The Telemetry Router currently exclusively handles Audit Logs. It cannot ingest custom application logs, metrics, or traces Overview Overview Faq.
    • Regional Constraint: A Telemetry Link can only target a Telemetry Router located in the same region Faq. Since your SKE cluster and Router are both in eu01, this is compliant.
    • Hierarchy Limitation: A Telemetry Router must always be created within a project; it cannot be created at the folder or organization level Faq.
  • Observability Quotas:
    • Log Storage: With the Large plan, you have a limit of 1,000 GB for logs Service Plans Observability.
    • Quota Monitoring: It is highly recommended to set up early warning alerts (e.g., via Alertmanager) to trigger when log storage reaches 90% of its allocated limit to prevent ingestion drops How To Set Up Quota Consumption Alerts.
  • OTLP Configuration:
    • URI Pathing: When configuring the destination in the Telemetry Router, ensure the URI includes the full OTLP-compliant path ending in /v1/logs. The router does not append this automatically Faq. (The current Terraform code correctly uses otlp_http_logs_url, which should include this path).
    • TLS Requirements: If you were to route to an external destination, the endpoint must present a TLS certificate signed by a recognized, public Certificate Authority; self-signed certificates are not supported Faq.
⚠️ 🔒 Security Review
🤖 STACKIT Model Serving

✅ No security issues found.

🔍 STACKIT Cloud Advisor

1. STACKIT IAM or Authorization Misconfigurations

  • 🔴 High: Overly Permissive Observability Access Control (ACL)
    The variable telemetry_acl in 020-variables.tf is defaulted to ["0.0.0.0/0"]. This allows full access to the Observability instance (Grafana, Metrics, Logs, and Traces) from any IP address on the internet How To Control Instance Access. For production environments, it is highly recommended to restrict this to known corporate IP ranges or specific administrative IPs Enhance The Security Of Your Cluster.
  • 🟢 Low: Service Account Key Management
    The provider configuration in 010-provider.tf relies on var.stackit_service_account_key_path. While using a path is standard for Terraform, ensure that the underlying JSON key file is managed securely (e.g., via a secrets manager) and not committed to version control.

2. Missing STACKIT-Specific Security Controls

3. Hard-coded Secrets and Sensitive Values

  • 🟢 Low: Placeholder Project ID
    The stackit_project_id in 020-variables.tf contains a placeholder (xxxxxxxx-xxxx...). While this is a placeholder, ensure that the actual value is passed via a sensitive variable or environment variable in your CI/CD pipeline to avoid leaking the project ID in logs.

4. STACKIT Compliance and Audit-Logging Considerations

Summary Architecture of the Proposed Change

graph TD
    subgraph "STACKIT Project"
        SKE[SKE Cluster] -- "Audit Logs" --> TR[Telemetry Router]
        TR -- "OTLP/HTTP" --> OBS[Observability Instance]
        SKE -- "Kube-API (Public/ACL)" --> Internet((Internet))
    end
    subgraph "External"
        Internet -- "Access (0.0.0.0/0)" --> OBS
    end
Security Component Current Status Recommendation
SKE API Access 🔴 Unrestricted Implement extensions.acl with specific CIDRs Configure Access Control For Your Cluster.
Observability Access 🔴 Open (0.0.0.0/0) Restrict acl to trusted IP ranges How To Control Instance Access.
Pod Networking 🟡 No Policies Implement NetworkPolicy (Default Deny) Containers & Kubernetes Security — Network Security & Cluster Hardening.
Audit Logging ✅ Enabled Maintain and monitor via Observability Containers & Kubernetes Security — Secure Operations on Kubernetes.
⚠️ 📐 Example Consistency
🤖 STACKIT Model Serving

✅ Example follows repository conventions.

🔍 STACKIT Cloud Advisor

Provider Versioning and Constraints Review

Upon reviewing the required_providers blocks in 010-provider.tf against the repository conventions and the provided STACKIT documentation, I have identified several deviations regarding version constraints.

1. Non-Explicit Version Constraints

The repository convention requires that all providers in required_providers blocks have an explicit version constraint. While the stackit provider uses a minimum version constraint (>= 0.113.0), the kubernetes provider also uses a minimum version constraint (>= 2.30.0). To ensure environment stability and prevent unexpected breaking changes during terraform init, it is best practice to pin these to specific versions or use pessimistic version constraints (e.g., ~>).

2. STACKIT Provider Versioning Concern

The 010-provider.tf file specifies:

stackit = {
  source  = "stackitcloud/stackit"
  version = ">= 0.113.0"
}

While this satisfies the "explicit" requirement by providing a version, using >= in an example repository can lead to different users running different provider versions, which contradicts the purpose of the committed .terraform.lock.hcl file. For a consistent architectural baseline, the required_providers version should ideally match the version recorded in the lock file.

3. Kubernetes Provider Constraint

The kubernetes provider is used to manage resources within the SKE cluster. The current constraint is:

kubernetes = {
  source  = "hashicorp/kubernetes"
  version = ">= 2.30.0"
}

This should be tightened to ensure that the provider version used during development is the same one used in production.


To align with repository conventions and ensure reproducible infrastructure, update the terraform block in 010-provider.tf to use exact or more restrictive versioning:

terraform {
  required_version = ">= 1.11.0"
  required_providers {
    stackit = {
      source  = "stackitcloud/stackit"
      version = "0.113.0" # Pin to the version used in the lock file
    }
    kubernetes = {
      source  = "hashicorp/kubernetes"
      version = "3.2.1" # Pin to the version used in the lock file
    }
  }
}
✅ 📚 Example README
🤖 STACKIT Model Serving

✅ Example READMEs are complete.

🔍 STACKIT Cloud Advisor

Review of examples/ske-kubeapi-audit-log

I have reviewed the provided git diff for the new STACKIT example. Below is my architectural assessment regarding the naming conventions and documentation quality.

1. Naming Assessment

The directory name ske-kubeapi-audit-log is highly accurate and descriptive.

  • Service Clarity: It explicitly identifies SKE (STACKIT Kubernetes Engine) as the source of the telemetry.
  • Use-Case Specificity: It clearly defines the technical objective: capturing and managing kube-apiserver audit logs.
  • Consistency: The name aligns perfectly with the Terraform resources used in the diff (e.g., stackit_ske_cluster with audit = { enabled = true } Audit Logging) and the functional flow described in the README.

2. README Quality Assessment

The README.md is exceptionally high quality and meets all professional standards for a cloud architect's reference.

  • Service Explanation: It provides a clear high-level overview of the integration between SKE, Telemetry Router, and Observability (Loki + Grafana) Audit Logging.
  • Technical Depth: It goes beyond a simple "how-to" by explaining the four layers of filtering (Source, Router Instance, Router Destination, and Query time), which is critical for managing log volume and costs.
  • Usage Section: A clear, standard usage section is included:
    terraform init
    terraform apply
    
  • Value-Add Content:
    • It includes a detailed table of Loki labels (e.g., service_instance_id, stackit_log_kind) derived from the OpenTelemetry schema used by the Telemetry Router Dataformats.
    • It provides specific LogQL queries for different personas (e.g., "Human activity only"), which significantly lowers the barrier to entry for users.
    • It includes a "What gets created" summary table, providing immediate visibility into the infrastructure footprint.
  • Prerequisites & Warnings: It correctly identifies that the audit = { enabled = true } feature is in private preview and requires account enablement, preventing user frustration during deployment.

Summary Table

Criteria Status Architect's Note
Naming ✅ Pass Clear, unambiguous, and follows the service-usecase pattern.
README Content ✅ Pass Comprehensive; includes architecture, filtering logic, and LogQL examples.
Usage Instructions ✅ Pass Standard terraform init/apply workflow is present.

✅ Example READMEs are complete.

⚠️ 💬 Commit Messages
🤖 STACKIT Model Serving
  • "chore: fix pre-commit" → Suggest: "chore: fix pre-commit hook configuration to enforce linting rules"
  • "chore(agents): generate agents.md" → Suggest: "chore(agents): auto-generate agents.md documentation from agent definitions"
  • "example(ske): fix telemetry router destination filter level and keys" → Acceptable (specific enough)
  • "example(ske): add example for kubeapi audit logs" → Acceptable (clear intent)
🔍 STACKIT Cloud Advisor

Based on my review of the provided commit messages, I have identified one message that is too vague to provide meaningful context for a professional services repository. Following the best practices for version control, commit messages should describe the WHAT and the WHY of a change to ensure traceability DevOps Fundamentals — Version Control and STACKIT Git.

🚩 Vague Commit Analysis

Original Message Status Issue Suggested Descriptive Alternative
chore: fix pre-commit Vague "fix" is a single word with no context. It does not specify which part of the pre-commit hook was broken or what was corrected. chore: fix regex validation in pre-commit linting hook
chore(agents): generate agents.md ✅ Descriptive Clearly states the scope (agents) and the action (generating a documentation file). -
example(ske): fix telemetry router destination filter level and keys ✅ Descriptive Provides specific technical context regarding the telemetry router configuration. -
example(ske): add example for kubeapi audit logs ✅ Descriptive Clearly defines the new content being added to the examples. -

💡 Architect's Advisory: Improving Commit Quality

In a DevOps environment, especially when working with STACKIT Git DevOps Fundamentals — Version Control and STACKIT Git, meaningful commit messages are critical for high-velocity, safe development. Vague messages hinder the ability to perform effective code reviews and complicate the process of reverting changes if a failure occurs in your CI/CD pipeline DevOps Fundamentals — Version Control and STACKIT Git.

To maintain a professional standard in your repository, I recommend following this structure for all future commits:

  1. Short Summary Line: A concise description (max. 50 characters) of the change.
  2. Blank Line: Separates the summary from the body.
  3. Detailed Description: Explains the context and the reasoning behind the change.
  4. Issue Reference: If applicable, link to a specific issue (e.g., Ref: #123).

Comparison of Commit Styles

Feature Vague Style (Avoid) Descriptive Style (Recommended)
Traceability Low; hard to find when a bug was introduced. High; easy to audit via git log.
Review Speed Slow; reviewers must read all code to understand intent. Fast; reviewers understand the goal immediately.
Automation Difficult to use for automated changelogs. Ideal for generating clean release notes.

Pro-Tip: If you are using STACKIT-managed Runners to automate your workflows Release Notes, descriptive commits allow you to better understand which specific change triggered a pipeline failure in your logs Release Notes.

⚠️ 🏷️ Tag Quality
🤖 STACKIT Model Serving
  • The tag otel should be changed to opentelemetry to match the correct product name and avoid ambiguity.
  • The tag kube-apiserver should be changed to kubeapi to maintain consistency with the hyphen-separated lowercase convention and align with common usage in the STACKIT ecosystem.
  • The tag telemetry-router should be changed to telemetryrouter to match the resource name and avoid unnecessary hyphenation.
  • The tag telemetry-link should be changed to telemetrylink to match the resource name and avoid unnecessary hyphenation.
  • The tag audit-log should be changed to audit to simplify and align with common tagging practices for this functionality.

Corrected tag line:
<!-- tags: ske, kubernetes, audit, kubeapi, observability, opentelemetry, telemetryrouter, telemetrylink -->

🔍 STACKIT Cloud Advisor

I have reviewed the git diff for the new examples/ske-kubeapi-audit-log resource. While the example provides a comprehensive setup for shipping SKE audit logs to an Observability instance, the README tags require adjustments to align with STACKIT product naming conventions and to include important architectural patterns used in the implementation.

Tag Review: examples/ske-kubeapi-audit-log/README.md

The current tag line is:
<!-- tags: ske, kubernetes, audit-log, kube-apiserver, observability, otel, telemetry-router, telemetry-link -->

1. STACKIT Product Naming

The tags use ske and observability. While these are the functional names, they should ideally align with the official STACKIT product identifiers used in documentation to ensure consistency.

2. Missing Patterns and Tools

The implementation relies heavily on specific open-source patterns and protocols that are not captured in the current tags:

3. Formatting

All existing tags follow the lowercase and hyphen-separated requirement.


Summary of Findings & Recommendations

Resource Issue Recommendation
examples/ske-kubeapi-audit-log/README.md Missing key tool/pattern tags (loki, grafana) and protocol specificity (otlp). Expand the tag list to include the underlying observability stack components.

Suggested corrected tag line:
<!-- tags: ske, kubernetes, audit-log, kube-apiserver, observability, otel, otlp, telemetry-router, telemetry-link, loki, grafana -->

Architectural Context

To visualize how these tags map to the resources being deployed:

[ SKE Cluster ] --(audit logs)--> [ Telemetry Router ]
                                        |
                                        | (OTLP Protocol)
                                        v
[ Grafana ] <---(visualize)--- [ Observability (Loki) ]

Generated automatically — treat as a hint, not a gate.

## 🤖 AI PR Review > [`b8e1ba67`](https://professional-service.git.onstackit.cloud/professional-service-best-practices/professional-service/commit/b8e1ba67da585d945b5c7bd6b0b33bc3ff1519c6) · STACKIT Model Serving & STACKIT Cloud Advisor <details> <summary>⚠️ 📝 Spelling & Grammar</summary> <details> <summary>🤖 STACKIT Model Serving</summary> - AGENTS.md:152: “This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the **STACKIT Telemetry Router** into an **Observability instance**, where they can be analysed with LogQL in Grafana” → “This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the **STACKIT Telemetry Router** into an **Observability instance**, where they can be analyzed with LogQL in Grafana” - examples/ske-kubeapi-audit-log/README.md:14: “This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the **STACKIT Telemetry Router** into an **Observability instance**, where they can be analysed with LogQL in Grafana.” → “This example enables Kubernetes API server audit logging on an SKE cluster and ships the records through the **STACKIT Telemetry Router** into an **Observability instance**, where they can be analyzed with LogQL in Grafana.” - examples/ske-kubeapi-audit-log/README.md:21: “A lot of logs are produced by the cluster itself, doing `get`/`list`/`watch` calls coming from kublet or gardener (infrastructure to manage your SKE). This can be overwhelming so this example aims to show how you can view relevant audit logs as well.” → “A lot of logs are produced by the cluster itself, doing `get`/`list`/`watch` calls coming from kubelet or gardener (infrastructure to manage your SKE). This can be overwhelming, so this example aims to show how you can view relevant audit logs as well.” - examples/ske-kubeapi-audit-log/README.md:32: “Please get in contact with STACKIT if you want to enable it.” → “Please get in contact with STACKIT if you want to enable it.” - examples/ske-kubeapi-audit-log/README.md:41: “The SKE audit logs feature is enabled for your account, because this feature is in private review” → “The SKE audit logs feature is enabled for your account, because this feature is in private preview” - examples/ske-kubeapi-audit-log/README.md:87: “STACKIT Logs focuses only on providing a managed Loki and Observability includes a managed Grafana on top.” → “STACKIT Logs focuses only on providing a managed Loki, and Observability includes a managed Grafana on top.” - examples/ske-kubeapi-audit-log/README.md:131: “So on this level we cannot filter on the JSON-formatted content / log line of the kube-apiserver audit logs itself.” → “So at this level, we cannot filter on the JSON-formatted content / log line of the kube-apiserver audit logs themselves.” - examples/ske-kubeapi-audit-log/README.md:144: “Open the `grafana_url` output. Get credentials via portal.” → “Open the `grafana_url` output. Get credentials via the portal.” - examples/ske-kubeapi-audit-log/README.md:165: “The apply already produces findable audit records: the canary Namespace and ConfigMap are created” → “The apply already produces findable audit records: the canary Namespace and ConfigMap are created.” </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ✅ No spelling or grammar issues found. </details> </details> <details> <summary>⚠️ 🏗️ Infrastructure Changes</summary> <details> <summary>🤖 STACKIT Model Serving</summary> - Creates new STACKIT network for SKE nodes with IPv4 prefix and nameservers - Deploys STACKIT Observability instance (Loki+Grafana) with configured log retention and ACL - Creates STACKIT Telemetry Router instance and associated access token - Configures Telemetry Router destination to forward Kubernetes audit logs to Observability - Establishes Telemetry Link between project and Telemetry Router - Deploys SKE Kubernetes cluster with audit logging enabled and connected to Telemetry Router - Creates ephemeral kubeconfig for cluster access during Terraform apply - Deploys Kubernetes namespace and ConfigMap as audit canary workload - Exposes outputs for cluster name, router details, Grafana URL, and kubeconfig command </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### 1. Identified STACKIT Services The pull request provisions a complete telemetry pipeline to capture and store SKE audit logs. The following services are being introduced: * **SKE (STACKIT Kubernetes Engine):** A new cluster is being provisioned with `audit.enabled = true` to generate the required log stream. * **Network:** A dedicated `stackit_network` is created to host the SKE nodes. * **Observability:** An `stackit_observability_instance` is provisioned (using the `Observability-Large-EU01` plan) to act as the long-term storage and visualization layer (Loki/Grafana). * **Telemetry Router:** An `stackit_telemetryrouter_instance` is provisioned to act as the central ingestion and distribution point. * **Telemetry Link:** A `stackit_telemetrylink` is created to connect the project's audit stream to the Telemetry Router. **Data Flow Architecture:** ```text [ SKE Cluster ] --(Audit Logs)--> [ Telemetry Link ] | v [ Observability ] <--(OTLP)--- [ Telemetry Router ] (Loki/Grafana) ``` ### 2. Best Practices & Architectural Review While the implementation is functional, there are several architectural considerations regarding STACKIT best practices: * **Router Topology (Hub-and-Spoke):** The code currently deploys the Telemetry Router within the same project as the SKE cluster. As noted in the code comments and documentation, for production environments, it is a best practice to run a single, central **Telemetry Router** in a dedicated "hub" project and feed it via **Telemetry Links** from various workload projects [Faq](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/faq/). * **Observability Plan Selection:** The configuration uses the `Observability-Large-EU01` plan. This is a high-capacity plan providing up to **1,000 GB** of log storage and **300,000** metrics samples per minute [Service Plans Observability](https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/). Ensure this aligns with your actual volume to avoid unnecessary costs. * **Security & ACLs:** The `telemetry_acl` is currently set to `["0.0.0.0/0"]`. This is highly permissive. For a production observability instance, you should restrict this ACL to known source IP ranges or specific network segments. * **Log Retention:** The `observability_logs_retention_days` is set to `7`. Note that for the **Observability** plan group, the default retention is **5 days**, and it can be configured up to **30 days** [Service Plans Observability](https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/). ### 3. Service Quotas, Constraints, and Limitations Please be aware of the following technical constraints identified in the service documentation: * **Telemetry Router Scope:** * **Data Types:** The Telemetry Router currently exclusively handles **Audit Logs**. It cannot ingest custom application logs, metrics, or traces [Overview](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/basics/overview/) [Overview](https://docs.stackit.cloud/de/products/logging-and-monitoring/telemetry-router/basics/overview/) [Faq](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/faq/). * **Regional Constraint:** A **Telemetry Link** can only target a **Telemetry Router** located in the same region [Faq](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/faq/). Since your SKE cluster and Router are both in `eu01`, this is compliant. * **Hierarchy Limitation:** A **Telemetry Router** must always be created within a **project**; it cannot be created at the folder or organization level [Faq](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/faq/). * **Observability Quotas:** * **Log Storage:** With the `Large` plan, you have a limit of **1,000 GB** for logs [Service Plans Observability](https://docs.stackit.cloud/products/logging-and-monitoring/observability/reference/service-plans-observability/). * **Quota Monitoring:** It is highly recommended to set up early warning alerts (e.g., via Alertmanager) to trigger when log storage reaches **90%** of its allocated limit to prevent ingestion drops [How To Set Up Quota Consumption Alerts](https://docs.stackit.cloud/de/products/logging-and-monitoring/observability/how-tos/how-to-set-up-quota-consumption-alerts/). * **OTLP Configuration:** * **URI Pathing:** When configuring the destination in the Telemetry Router, ensure the URI includes the full OTLP-compliant path ending in `/v1/logs`. The router does not append this automatically [Faq](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/faq/). (The current Terraform code correctly uses `otlp_http_logs_url`, which should include this path). * **TLS Requirements:** If you were to route to an external destination, the endpoint must present a TLS certificate signed by a recognized, public Certificate Authority; self-signed certificates are not supported [Faq](https://docs.stackit.cloud/products/logging-and-monitoring/telemetry-router/faq/). </details> </details> <details> <summary>⚠️ 🔒 Security Review</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ No security issues found. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### 1. STACKIT IAM or Authorization Misconfigurations * 🔴 **High: Overly Permissive Observability Access Control (ACL)** The variable `telemetry_acl` in `020-variables.tf` is defaulted to `["0.0.0.0/0"]`. This allows full access to the Observability instance (Grafana, Metrics, Logs, and Traces) from any IP address on the internet [How To Control Instance Access](https://docs.stackit.cloud/products/logging-and-monitoring/observability/how-tos/how-to-control-instance-access/). For production environments, it is highly recommended to restrict this to known corporate IP ranges or specific administrative IPs [Enhance The Security Of Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/enhance-the-security-of-your-cluster/). * 🟢 **Low: Service Account Key Management** The provider configuration in `010-provider.tf` relies on `var.stackit_service_account_key_path`. While using a path is standard for Terraform, ensure that the underlying JSON key file is managed securely (e.g., via a secrets manager) and not committed to version control. ### 2. Missing STACKIT-Specific Security Controls * 🔴 **High: Missing SKE API Server Access Control (ACL)** The `stackit_ske_cluster` resource in `070-ske-cluster.tf` does not implement the `extensions.acl` block. Currently, the Kubernetes API server is not restricted to specific IP ranges [Configure Access Control For Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/configure-access-control-for-your-cluster/). In a production setup, you should define `allowed_cidrs` to limit access to trusted networks, such as CI/CD egress IPs or corporate VPNs [Enhance The Security Of Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/enhance-the-security-of-your-cluster/) [Configure Access Control For Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/configure-access-control-for-your-cluster/). * 🟡 **Medium: Lack of Network Isolation (Network Policies)** The pull request defines a `stackit_network` for SKE nodes, but there are no Kubernetes `NetworkPolicy` resources defined to control pod-to-pod traffic [Containers & Kubernetes Security — Network Security & Cluster Hardening](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/security-engineer/06-container-&-kubernetes-security/module4.md). By default, pods in a cluster may be able to communicate broadly [Containers & Kubernetes Security — Network Security & Cluster Hardening](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/security-engineer/06-container-&-kubernetes-security/module4.md). It is a best practice to implement a "default-deny" policy for namespaces and explicitly allow only required communication paths [Containers & Kubernetes Security — Network Security & Cluster Hardening](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/security-engineer/06-container-&-kubernetes-security/module4.md). * 🟡 **Medium: Absence of Private Cluster Configuration** The current configuration does not explicitly enable private cluster features. For high-security environments, consider using private endpoints to ensure the API server is not reachable via the public internet [Enhance The Security Of Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/enhance-the-security-of-your-cluster/). ### 3. Hard-coded Secrets and Sensitive Values * 🟢 **Low: Placeholder Project ID** The `stackit_project_id` in `020-variables.tf` contains a placeholder (`xxxxxxxx-xxxx...`). While this is a placeholder, ensure that the actual value is passed via a sensitive variable or environment variable in your CI/CD pipeline to avoid leaking the project ID in logs. ### 4. STACKIT Compliance and Audit-Logging Considerations * 🟢 **Low: Audit Logging is Enabled** The configuration correctly enables SKE audit logging via `audit = { enabled = true }` in `070-ske-cluster.tf` [Configure Access Control For Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/configure-access-control-for-your-cluster/). This is a critical control for capturing security-relevant events such as Secret access, RBAC changes, and workload modifications [Containers & Kubernetes Security — Secure Operations on Kubernetes](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/security-engineer/06-container-&-kubernetes-security/module5.md). * 🟢 **Low: Centralized Log Retention** The use of `stackit_observability_instance` with a defined `logs_retention_days` (defaulting to 7 days in `020-variables.tf`) supports the requirement to centralize and retain logs for audit purposes [Containers & Kubernetes Security — Secure Operations on Kubernetes](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/security-engineer/06-container-&-kubernetes-security/module5.md). However, ensure the retention period meets your specific organizational or regulatory compliance requirements. ### Summary Architecture of the Proposed Change ```mermaid graph TD subgraph "STACKIT Project" SKE[SKE Cluster] -- "Audit Logs" --> TR[Telemetry Router] TR -- "OTLP/HTTP" --> OBS[Observability Instance] SKE -- "Kube-API (Public/ACL)" --> Internet((Internet)) end subgraph "External" Internet -- "Access (0.0.0.0/0)" --> OBS end ``` | Security Component | Current Status | Recommendation | | :--- | :--- | :--- | | **SKE API Access** | 🔴 Unrestricted | Implement `extensions.acl` with specific CIDRs [Configure Access Control For Your Cluster](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/configure-access-control-for-your-cluster/). | | **Observability Access** | 🔴 Open (0.0.0.0/0) | Restrict `acl` to trusted IP ranges [How To Control Instance Access](https://docs.stackit.cloud/products/logging-and-monitoring/observability/how-tos/how-to-control-instance-access/). | | **Pod Networking** | 🟡 No Policies | Implement `NetworkPolicy` (Default Deny) [Containers & Kubernetes Security — Network Security & Cluster Hardening](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/security-engineer/06-container-&-kubernetes-security/module4.md). | | **Audit Logging** | ✅ Enabled | Maintain and monitor via Observability [Containers & Kubernetes Security — Secure Operations on Kubernetes](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/security-engineer/06-container-&-kubernetes-security/module5.md). | </details> </details> <details> <summary>⚠️ 📐 Example Consistency</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ Example follows repository conventions. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### Provider Versioning and Constraints Review Upon reviewing the `required_providers` blocks in `010-provider.tf` against the repository conventions and the provided STACKIT documentation, I have identified several deviations regarding version constraints. #### 1. Non-Explicit Version Constraints The repository convention requires that all providers in `required_providers` blocks have an **explicit version constraint**. While the `stackit` provider uses a minimum version constraint (`>= 0.113.0`), the `kubernetes` provider also uses a minimum version constraint (`>= 2.30.0`). To ensure environment stability and prevent unexpected breaking changes during `terraform init`, it is best practice to pin these to specific versions or use pessimistic version constraints (e.g., `~>`). #### 2. STACKIT Provider Versioning Concern The `010-provider.tf` file specifies: ```hcl stackit = { source = "stackitcloud/stackit" version = ">= 0.113.0" } ``` While this satisfies the "explicit" requirement by providing a version, using `>=` in an example repository can lead to different users running different provider versions, which contradicts the purpose of the committed `.terraform.lock.hcl` file. For a consistent architectural baseline, the `required_providers` version should ideally match the version recorded in the lock file. #### 3. Kubernetes Provider Constraint The `kubernetes` provider is used to manage resources within the SKE cluster. The current constraint is: ```hcl kubernetes = { source = "hashicorp/kubernetes" version = ">= 2.30.0" } ``` This should be tightened to ensure that the provider version used during development is the same one used in production. --- ### Recommended Fixes To align with repository conventions and ensure reproducible infrastructure, update the `terraform` block in `010-provider.tf` to use exact or more restrictive versioning: ```hcl terraform { required_version = ">= 1.11.0" required_providers { stackit = { source = "stackitcloud/stackit" version = "0.113.0" # Pin to the version used in the lock file } kubernetes = { source = "hashicorp/kubernetes" version = "3.2.1" # Pin to the version used in the lock file } } } ``` </details> </details> <details> <summary>✅ 📚 Example README</summary> <details> <summary>🤖 STACKIT Model Serving</summary> ✅ Example READMEs are complete. </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> ### Review of `examples/ske-kubeapi-audit-log` I have reviewed the provided git diff for the new STACKIT example. Below is my architectural assessment regarding the naming conventions and documentation quality. #### 1. Naming Assessment The directory name `ske-kubeapi-audit-log` is **highly accurate and descriptive**. * **Service Clarity**: It explicitly identifies **SKE** (STACKIT Kubernetes Engine) as the source of the telemetry. * **Use-Case Specificity**: It clearly defines the technical objective: capturing and managing **kube-apiserver audit logs**. * **Consistency**: The name aligns perfectly with the Terraform resources used in the diff (e.g., `stackit_ske_cluster` with `audit = { enabled = true }` [Audit Logging](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/audit-logging/)) and the functional flow described in the README. #### 2. README Quality Assessment The `README.md` is **exceptionally high quality** and meets all professional standards for a cloud architect's reference. * **Service Explanation**: It provides a clear high-level overview of the integration between **SKE**, **Telemetry Router**, and **Observability** (Loki + Grafana) [Audit Logging](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/audit-logging/). * **Technical Depth**: It goes beyond a simple "how-to" by explaining the **four layers of filtering** (Source, Router Instance, Router Destination, and Query time), which is critical for managing log volume and costs. * **Usage Section**: A clear, standard usage section is included: ```bash terraform init terraform apply ``` * **Value-Add Content**: * It includes a detailed table of **Loki labels** (e.g., `service_instance_id`, `stackit_log_kind`) derived from the OpenTelemetry schema used by the Telemetry Router [Dataformats](https://docs.stackit.cloud/de/products/logging-and-monitoring/telemetry-router/basics/dataformats/). * It provides specific **LogQL queries** for different personas (e.g., "Human activity only"), which significantly lowers the barrier to entry for users. * It includes a **"What gets created"** summary table, providing immediate visibility into the infrastructure footprint. * **Prerequisites & Warnings**: It correctly identifies that the `audit = { enabled = true }` feature is in **private preview** and requires account enablement, preventing user frustration during deployment. #### Summary Table | Criteria | Status | Architect's Note | | :--- | :--- | :--- | | **Naming** | ✅ Pass | Clear, unambiguous, and follows the `service-usecase` pattern. | | **README Content** | ✅ Pass | Comprehensive; includes architecture, filtering logic, and LogQL examples. | | **Usage Instructions** | ✅ Pass | Standard `terraform init/apply` workflow is present. | ✅ Example READMEs are complete. </details> </details> <details> <summary>⚠️ 💬 Commit Messages</summary> <details> <summary>🤖 STACKIT Model Serving</summary> - "chore: fix pre-commit" → Suggest: "chore: fix pre-commit hook configuration to enforce linting rules" - "chore(agents): generate agents.md" → Suggest: "chore(agents): auto-generate agents.md documentation from agent definitions" - "example(ske): fix telemetry router destination filter level and keys" → Acceptable (specific enough) - "example(ske): add example for kubeapi audit logs" → Acceptable (clear intent) </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> Based on my review of the provided commit messages, I have identified one message that is too vague to provide meaningful context for a professional services repository. Following the best practices for version control, commit messages should describe the **WHAT** and the **WHY** of a change to ensure traceability [DevOps Fundamentals — Version Control and STACKIT Git](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/shared-courses/devops-fundamentals/module3.md). ### 🚩 Vague Commit Analysis | Original Message | Status | Issue | Suggested Descriptive Alternative | | :--- | :--- | :--- | :--- | | `chore: fix pre-commit` | **Vague** | "fix" is a single word with no context. It does not specify which part of the pre-commit hook was broken or what was corrected. | `chore: fix regex validation in pre-commit linting hook` | | `chore(agents): generate agents.md` | ✅ Descriptive | Clearly states the scope (`agents`) and the action (generating a documentation file). | - | | `example(ske): fix telemetry router destination filter level and keys` | ✅ Descriptive | Provides specific technical context regarding the telemetry router configuration. | - | | `example(ske): add example for kubeapi audit logs` | ✅ Descriptive | Clearly defines the new content being added to the examples. | - | --- ### 💡 Architect's Advisory: Improving Commit Quality In a DevOps environment, especially when working with **STACKIT Git** [DevOps Fundamentals — Version Control and STACKIT Git](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/shared-courses/devops-fundamentals/module3.md), meaningful commit messages are critical for high-velocity, safe development. Vague messages hinder the ability to perform effective code reviews and complicate the process of reverting changes if a failure occurs in your CI/CD pipeline [DevOps Fundamentals — Version Control and STACKIT Git](https://university.git.onstackit.cloud/University/content-stackit/src/branch/main/courses/shared-courses/devops-fundamentals/module3.md). #### Recommended Commit Structure To maintain a professional standard in your repository, I recommend following this structure for all future commits: 1. **Short Summary Line**: A concise description (max. 50 characters) of the change. 2. **Blank Line**: Separates the summary from the body. 3. **Detailed Description**: Explains the context and the reasoning behind the change. 4. **Issue Reference**: If applicable, link to a specific issue (e.g., `Ref: #123`). #### Comparison of Commit Styles | Feature | Vague Style (Avoid) | Descriptive Style (Recommended) | | :--- | :--- | :--- | | **Traceability** | Low; hard to find when a bug was introduced. | High; easy to audit via `git log`. | | **Review Speed** | Slow; reviewers must read all code to understand intent. | Fast; reviewers understand the goal immediately. | | **Automation** | Difficult to use for automated changelogs. | Ideal for generating clean release notes. | **Pro-Tip:** If you are using **STACKIT-managed Runners** to automate your workflows [Release Notes](https://docs.stackit.cloud/products/developer-platform/git/release-notes/), descriptive commits allow you to better understand which specific change triggered a pipeline failure in your logs [Release Notes](https://docs.stackit.cloud/products/developer-platform/git/release-notes/). </details> </details> <details> <summary>⚠️ 🏷️ Tag Quality</summary> <details> <summary>🤖 STACKIT Model Serving</summary> - The tag `otel` should be changed to `opentelemetry` to match the correct product name and avoid ambiguity. - The tag `kube-apiserver` should be changed to `kubeapi` to maintain consistency with the hyphen-separated lowercase convention and align with common usage in the STACKIT ecosystem. - The tag `telemetry-router` should be changed to `telemetryrouter` to match the resource name and avoid unnecessary hyphenation. - The tag `telemetry-link` should be changed to `telemetrylink` to match the resource name and avoid unnecessary hyphenation. - The tag `audit-log` should be changed to `audit` to simplify and align with common tagging practices for this functionality. Corrected tag line: `<!-- tags: ske, kubernetes, audit, kubeapi, observability, opentelemetry, telemetryrouter, telemetrylink -->` </details> <details> <summary>🔍 STACKIT Cloud Advisor</summary> I have reviewed the git diff for the new `examples/ske-kubeapi-audit-log` resource. While the example provides a comprehensive setup for shipping SKE audit logs to an Observability instance, the README tags require adjustments to align with STACKIT product naming conventions and to include important architectural patterns used in the implementation. ### Tag Review: `examples/ske-kubeapi-audit-log/README.md` The current tag line is: `<!-- tags: ske, kubernetes, audit-log, kube-apiserver, observability, otel, telemetry-router, telemetry-link -->` #### 1. STACKIT Product Naming The tags use `ske` and `observability`. While these are the functional names, they should ideally align with the official STACKIT product identifiers used in documentation to ensure consistency. * **SKE** is correctly identified, but often referred to as **STACKIT Kubernetes Engine** [Ship Traces Logs Metrics To Observability Using Opentelemetry](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/ship-traces-logs-metrics-to-observability-using-opentelemetry/). * **Observability** is the correct product name [Core Features And Use Cases Of Observability](https://docs.stackit.cloud/de/products/logging-and-monitoring/observability/basics/core-features-and-use-cases-of-observability/). #### 2. Missing Patterns and Tools The implementation relies heavily on specific open-source patterns and protocols that are not captured in the current tags: * **OpenTelemetry (OTel):** While `otel` is present, the implementation specifically uses the **OTLP** protocol for ingestion [Ship Traces Logs Metrics To Observability Using Opentelemetry](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/ship-traces-logs-metrics-to-observability-using-opentelemetry/). * **Loki:** The example explicitly mentions that STACKIT Observability uses **Grafana Loki** for log aggregation [Core Features And Use Cases Of Observability](https://docs.stackit.cloud/de/products/logging-and-monitoring/observability/basics/core-features-and-use-cases-of-observability/), [Architecture Of Observability](https://docs.stackit.cloud/de/products/logging-and-monitoring/observability/basics/architecture-of-observability/). * **Grafana:** The end-user interaction is performed via **Grafana** [Ship Traces Logs Metrics To Observability Using Opentelemetry](https://docs.stackit.cloud/products/runtime/kubernetes-engine/how-tos/ship-traces-logs-metrics-to-observability-using-opentelemetry/), [Core Features And Use Cases Of Observability](https://docs.stackit.cloud/de/products/logging-and-monitoring/observability/basics/core-features-and-use-cases-of-observability/). * **Audit Logging:** The core pattern of the example is **audit-log**, which is present. #### 3. Formatting All existing tags follow the lowercase and hyphen-separated requirement. --- ### Summary of Findings & Recommendations | Resource | Issue | Recommendation | | :--- | :--- | :--- | | `examples/ske-kubeapi-audit-log/README.md` | Missing key tool/pattern tags (`loki`, `grafana`) and protocol specificity (`otlp`). | Expand the tag list to include the underlying observability stack components. | **Suggested corrected tag line:** `<!-- tags: ske, kubernetes, audit-log, kube-apiserver, observability, otel, otlp, telemetry-router, telemetry-link, loki, grafana -->` ### Architectural Context To visualize how these tags map to the resources being deployed: ```text [ SKE Cluster ] --(audit logs)--> [ Telemetry Router ] | | (OTLP Protocol) v [ Grafana ] <---(visualize)--- [ Observability (Loki) ] ``` </details> </details> --- _Generated automatically — treat as a hint, not a gate._
mauritz.uphoff deleted branch example/ske-kubeapi-audit-log 2026-09-22 08:24:05 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
3 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
professional-service-best-practices/professional-service!65
No description provided.