# Clouds of Europe — Full Content Export > Reclaiming European digital independence through cloud sovereignty and Kubernetes Site: https://clouds-of-europe.eu Total articles: 58 Generated: 2026-09-05T07:56:56.021Z --- ## NIS2 Personal Liability: What Management Needs to Know URL: https://clouds-of-europe.eu/content/policy-sovereignty/policy-regulation/nis2-personal-liability-what-management-needs-to-know Author: Jurg van Vliet Published: 2026-04-07 Category: Policy & Sovereignty Type: Policy & Regulation There is a sentence in NIS2 that most board members have not read. It is in Article 20(1): > "Member States shall ensure that the management bodies of essential and important entities approve the cybersecurity risk-management measures taken by those entities in order to comply with Article 21, oversee its implementation and **can be held liable** for infringements by the entities of that Article." Three obligations in one sentence. You must approve the cybersecurity measures. You must oversee their implementation. And you can be held personally liable if they are inadequate. This is not a theoretical risk buried in a legal annex. It is the central governance provision of the most significant cybersecurity regulation Europe has produced. And the enforcement apparatus is now operational. ## What "personally liable" means in practice The directive sets a floor. Member states set the ceiling. And some have set it high. **Administrative liability** exists everywhere. Regulators can impose fines on the organisation — up to EUR 10 million or 2% of global annual turnover for essential entities, EUR 7 million or 1.4% for important entities. But those fines hit the company, not you personally. What hits you personally is Article 32(5). For essential entities, when other enforcement measures have failed, authorities can request that a court **temporarily prohibit any natural person responsible for discharging managerial responsibilities at chief executive officer or legal representative level from exercising managerial functions**. The prohibition lasts until the entity remedies its deficiencies. There is no fixed term. Read that again. If your organisation fails to comply and does not fix it after being told to, you can be barred from your role — indefinitely — until compliance is achieved. **Italy goes further.** The Italian transposition (Decreto Legislativo 138/2024) extends the management ban to both essential and important entities — beyond what the directive requires. And under Italian law, directors hold a "guarantee position" (posizione di garanzia) under criminal law. If a cyber incident causes criminally relevant consequences, directors can face criminal prosecution for failure to oversee NIS2 compliance. Delegation of operational tasks does not eliminate this criminal responsibility. **Germany codified civil liability.** Section 38(2) of the new BSI Act (BSIG) makes management body members personally liable for damages incurred from breach of their cybersecurity oversight duty. The company can bring internal claims against its own directors. An earlier draft prohibited the company from waiving these claims — that prohibition was deleted from the final version, but the underlying personal liability remains. **The Netherlands** is still finalising its transposition (the Cyberbeveiligingswet). The draft introduces explicit personal board liability, with boards required to demonstrate approval of cybersecurity measures. Board members must acquire sufficient cybersecurity knowledge within two years of the law entering force. ## The training requirement you cannot delegate Article 20(2) mandates that management body members undergo cybersecurity training. This is not a suggestion. It is a legal requirement. The purpose is specific: management must gain "sufficient knowledge and skills to enable them to identify risks and assess cybersecurity risk-management practices and their impact on the services provided by the entity." Germany requires this training at least every three years, with documentation of participants, content, trainers, and duration — available for potential audits. The explanatory memorandum states that "a simple certificate of attendance will not suffice." The training obligation is directly tied to the liability clause. A management body cannot credibly claim it approved adequate cybersecurity measures if its members were never trained to understand cybersecurity risks. If you signed off on a cybersecurity programme you did not understand, the training requirement makes that a documented governance failure. ## The ten measures you must have approved Article 21(2) lists ten minimum cybersecurity risk-management measures. All ten are mandatory. They are not a menu to choose from: 1. Policies on risk analysis and information system security 2. Incident handling 3. Business continuity, backup management, disaster recovery, and crisis management 4. **Supply chain security** — including security-related aspects of relationships with direct suppliers and service providers 5. Security in network and information systems acquisition, development, and maintenance 6. Policies and procedures to assess the effectiveness of cybersecurity measures 7. Basic cyber hygiene practices and cybersecurity training 8. Policies and procedures regarding cryptography and encryption 9. Human resources security, access control, and asset management 10. Multi-factor authentication, secured communications Article 20(1) says management can be held liable for infringements of Article 21. Article 21 makes all ten measures mandatory. If your organisation's cybersecurity programme is missing any of them — supply chain security, incident handling, business continuity — you have approved measures that do not comply with the directive. And you are personally exposed. The supply chain provision (number 4) is particularly relevant after the axios npm compromise of March 2026, which demonstrated that a three-hour window on a compromised package can affect organisations globally. If your board approved cybersecurity measures that did not address software supply chain risk, that is now a documented gap against an explicit regulatory requirement. ## What a NIS2 incident actually costs you When a significant incident occurs, NIS2 requires three reports: - **24 hours:** Early warning to the national authority. Must include whether the incident is suspected to be caused by unlawful or malicious acts and whether it could have cross-border impact. - **72 hours:** Full incident notification with initial severity assessment, indicators of compromise, and affected services. - **30 days:** Final report with root cause analysis, mitigation measures, and cross-border impact assessment. Missing these deadlines is independently sanctionable — the same penalty framework applies. Under NIS1, the Dutch authorities fined a telecommunications provider EUR 525,000 for failing to report a significant incident in a timely manner. NIS2 penalties are an order of magnitude higher. The practical cost of an incident without adequate preparation: 48+ hours of senior engineering time in a war room, legal counsel engaged on reporting obligations, potential customer notification, regulatory communication — and the clock starts the moment you detect the incident. If your organisation cannot answer "which of our services is affected and who owns them" within hours, you are already behind on the 24-hour reporting deadline. Authorities have signalled tolerance for early, honest imperfection — incomplete reports filed on time are treated more favourably than late or absent reports. But systematic failure to meet reporting obligations constitutes evidence of inadequate risk management, which compounds the personal liability exposure under Article 20. ## Your D&O insurance may not cover this Directors and Officers insurance policies were not designed for cyber governance failures. The gap is significant. D&O policies typically exclude claims arising from data breaches, regulatory fines, and third-party liabilities from cyber incidents. Cyber insurance policies cover the entity's losses from cyber events but do not cover personal director liability. NIS2 creates a scenario where neither policy clearly covers the personal claim. The additional problem: D&O policies rarely cover gross negligence or wilful neglect. If a director knew about cybersecurity deficiencies and failed to act, the insurer may deny the claim. NIS2's explicit governance requirements — approve, oversee, train — make it very difficult to argue ignorance. If you were trained, if you approved the measures, and if those measures were inadequate, the insurer's position is straightforward. Allianz's 2026 D&O Insurance Insights report identifies that claims against directors are increasingly triggered by data breaches, ransomware attacks, and technical failures. Ransomware accounted for approximately 60% of the value of large cyber insurance claims (greater than EUR 1 million) in the first half of 2025. The practical step: review your D&O policy with your broker. Ask whether it affirmatively covers NIS2 regulatory proceedings and personal liability arising from cyber governance failures. If cyber exclusions exist, they need to be addressed before the first audit. ## The GDPR precedent If you think personal liability under EU regulation is theoretical, consider the Clearview AI case. In September 2024, the Dutch Data Protection Authority fined Clearview AI EUR 30.5 million for illegal facial image scraping. The DPA then announced it was investigating whether to hold Clearview AI's directors personally liable — a first in GDPR enforcement history. The DPA's chairman stated: "This liability already exists if directors know that the GDPR is being violated, have the authority to stop that, but omit to do so, and in this way consciously accept those violations." GDPR does not contain explicit personal liability language comparable to NIS2 Article 20. If regulators are already pursuing individual directors under GDPR's implicit framework, NIS2's explicit personal liability provisions will be enforced with greater force and clearer legal basis. ## What to do about it This is not a checklist for compliance consultants. It is a governance question for anyone whose name is on the management body. **Verify your cybersecurity programme covers all ten Article 21 measures.** If you cannot confirm that supply chain security, incident handling, and business continuity are explicitly addressed, you have an immediate gap. Each missing measure is a potential infringement for which you are personally liable. **Confirm you have formally approved the measures.** The directive requires management body approval — not delegation to a CISO. If the cybersecurity programme was never formally presented to and approved by the management body, the approval requirement of Article 20 has not been met. **Complete cybersecurity training.** If you have not undergone training that enables you to "identify risks and assess cybersecurity risk-management practices," you are in violation of Article 20(2). Germany requires this every three years with substantive documentation. **Test your incident response capability.** Can your organisation answer "which services are affected and who owns them" within hours of a supply chain advisory? Can you produce the early warning report within 24 hours? If not, the incident reporting obligations of Article 23 are at risk, and the failure compounds your Article 20 exposure. **Review your D&O policy.** Confirm it covers NIS2 regulatory proceedings and personal liability for cyber governance failures. If it does not, address this with your broker before the first audit. **Document everything.** The difference between personal liability and a defensible governance position is documentation. Board minutes showing approval, training records with substantive content, incident response test results, supply chain risk assessments — these are the artefacts that demonstrate you took your oversight obligation seriously. The first NIS2 compliance audits are underway. The question is not whether the personal liability provisions will be tested — it is which organisation and which management body will be first. --- *This article is published on [Clouds of Europe](https://clouds-of-europe.eu), a practitioner community building European cloud independence. We assess regulatory impact based on directive text and national transposition law — no vendor sponsorship, no product agenda.* --- ## Your IDP Is Your Supply Chain Defence URL: https://clouds-of-europe.eu/content/practice/implementation-patterns/your-idp-is-your-supply-chain-defence Author: Jurg van Vliet Published: 2026-04-01 Category: Practice Type: Implementation Patterns At 00:21 UTC on March 31, a new version of axios — the HTTP client library used by 174,000 packages and downloaded 101 million times a week — was published to npm. Version 1.14.1 looked routine. It was not. A compromised maintainer account published the release directly, bypassing the project's normal CI/CD pipeline. The legitimate 1.14.0 release had used GitHub Actions with cryptographic publishing verification. The attacker skipped all of that and published straight from the hijacked account. The malicious version included a dependency called plain-crypto-js, which ran an install script that downloaded malware onto the build machine. That malware gave the attacker remote access: file browsing, command execution, and a persistent callback to their server every 60 seconds. Google Threat Intelligence attributed the attack to a North Korean group designated UNC1069. Security firm Huntress confirmed 135 endpoints contacted the attacker's server during the three-hour window before npm pulled the package. Elastic Security Labs found a 3% execution rate across affected environments — meaning in most cases, the malicious code reached machines but something in the build or runtime environment prevented full execution. Socket, a supply chain security monitoring service, detected the malware within six minutes of publication. npm removed the package roughly three hours later. For the organisations that pulled axios during that window, three hours was all it took. ## Why this keeps happening The axios compromise is not an anomaly. It is a recurring pattern in software supply chain attacks, and the pattern has been consistent for nearly a decade. In 2018, the event-stream package was compromised through a social engineering handover — a new maintainer injected code targeting a cryptocurrency wallet. In 2021, ua-parser-js, with 8 million weekly downloads, was hijacked to install cryptominers. In 2022, the maintainer of colors.js and faker.js deliberately corrupted his own packages in protest, breaking thousands of builds. The same year, the node-ipc package was weaponised with code targeting Russian IP addresses during the Ukraine war. In 2024, the xz-utils backdoor showed this pattern extending beyond npm into core Linux infrastructure, with a multi-year social engineering campaign to gain maintainer trust. Sonatype's 2026 report counted 454,648 malicious packages published to open source registries in 2025 alone — a 75% increase year over year. Cybersecurity Ventures estimated the global cost of software supply chain attacks at $60 billion in 2025, projected to reach $138 billion by 2031. The structural problem is straightforward. Public package registries are centralised, open-publish systems. Anyone with an account can publish. Packages pull in other packages transitively — your application may depend on hundreds of libraries you never explicitly chose. There is no verification layer between the public registry and your build system unless you build one. This is not a flaw in npm specifically. It is the architecture of how open source distribution works across every language ecosystem. The axios incident exploited it. The next incident will too. ## What your IDP should do An Internal Developer Platform that takes supply chain security seriously needs four capabilities. None of them require commercial tooling. All of them require deliberate engineering. ### Package registry as proxy The single most effective supply chain defence is not running builds directly against public registries. A self-hosted package registry that proxies npm, PyPI, Maven Central, and other upstream sources gives you a cache of known-good package versions and a control point during incidents. GitLab Community Edition includes a package registry that supports npm, PyPI, Maven, NuGet, and several other formats — at no licence cost. When configured as a proxy, it caches packages on first download and serves subsequent requests from the local cache. During the axios incident, organisations with a proxying registry had options. They could freeze their registry to prevent new versions from being pulled. They could audit exactly which projects had downloaded the compromised version. They could replace the cached malicious version with the known-good 1.14.0 and resume builds — all before npm completed their own cleanup. Organisations pulling directly from npm had one option: wait. ### CI-enforced defences The axios attack vector was an install script — code that runs automatically when npm installs a package. This is a known and well-documented attack surface. The defence is a single CI pipeline flag: `--ignore-scripts`. Running `npm install --ignore-scripts` (or the equivalent `yarn install --ignore-scripts`) disables all lifecycle scripts during installation. It blocks the exact mechanism the axios malware used. Most applications work fine without install scripts. The ones that don't — typically packages with native binary compilation — need explicit exceptions, which forces you to enumerate and review the packages that have this elevated privilege. Beyond script blocking, your CI pipeline should enforce lockfile integrity. Every build should verify that the installed dependency tree matches the committed lockfile exactly. If someone updates a dependency without updating the lockfile, the build should fail. This catches both supply chain attacks and accidental dependency drift. Signature verification is the next layer. npm now supports package provenance — cryptographic attestation that a package was built by a specific CI system from a specific source repository. The legitimate axios 1.14.0 used this. The malicious 1.14.1 did not, because it was published directly from a compromised account rather than through CI. Verifying provenance in your pipeline would have flagged this discrepancy automatically. ### Dependency scanning and SBOM generation You need to know what you depend on, whether any of it has known vulnerabilities, and you need that information in a machine-readable format. Three open source tools handle this: - **Trivy** (Aqua Security, Apache 2.0 licence) scans container images, filesystems, and Git repositories for known vulnerabilities. It covers OS packages, language-specific dependencies, and infrastructure-as-code misconfigurations. - **Syft** (Anchore, Apache 2.0 licence) generates Software Bills of Materials in standard formats (SPDX, CycloneDX) from container images and filesystems. - **osv-scanner** (Google, Apache 2.0 licence) checks dependencies against the OSV database — the largest aggregated vulnerability database, covering multiple ecosystems. These are not hobbyist tools. AWS uses Trivy in several internal workflows. Google maintains osv-scanner and the OSV database behind it. They are production-grade and free. One important caveat: 65% of open source vulnerabilities lack a severity score in the National Vulnerability Database, with a median 41-day lag between disclosure and scoring. Tools that rely solely on NVD miss a significant portion of the risk. osv-scanner uses the broader OSV database precisely because of this coverage gap. GitLab Ultimate includes native dependency scanning and SBOM features, but GitLab CE combined with these three open source tools achieves the same outcome at zero licence cost. The tools run as CI pipeline stages and produce artefacts that your registry can store alongside your build output. ### Impact assessment through your service catalogue Knowing that axios is compromised is step one. Knowing which of your services use axios, who owns those services, which ones are customer-facing, and which ones run in production — that is step two, and it is where most organisations fumble during an incident. Backstage, configured as your service catalogue, makes this queryable. When a vulnerability is disclosed, you need to answer: "which services are affected and what is the blast radius?" A well-maintained catalogue with dependency metadata lets you answer that question in minutes rather than hours. The Backstage SBOM plugin can aggregate Software Bills of Materials across your entire service catalogue, making organisation-wide dependency queries possible. "Show me every service that depends on axios" becomes a catalogue query rather than a Slack thread asking team leads to check their package.json files. This is the difference between incident response and incident archaeology. One happens in real time. The other happens over the next week. ## The regulatory floor For European organisations, supply chain security is no longer a best practice recommendation. It is a legal obligation. NIS2 Article 21(2)(d) explicitly requires "supply chain security, including security-related aspects concerning the relationships between each entity and its direct suppliers or service providers." The directive's first compliance audits begin June 30, 2026. Penalties for non-compliance reach EUR 10 million or 2% of global annual turnover, whichever is higher. The Cyber Resilience Act adds another layer. From September 11, 2026, products with digital elements sold in the EU must comply with SBOM reporting obligations. If your software includes open source dependencies — and it does — you need to produce and maintain Software Bills of Materials as a regulatory requirement. These are not abstract future risks. The NIS2 audit deadline is less than three months away. Organisations that cannot demonstrate supply chain security controls — a registry policy, CI-enforced scanning, SBOM generation, and an incident response process for compromised dependencies — will have a gap in their compliance posture that auditors are specifically trained to find. ## The sovereignty argument There is a sovereignty dimension to this that goes beyond compliance. The npm registry is a centralised system operated by GitHub, which is owned by Microsoft, which is a US-headquartered corporation subject to US jurisdiction. Every `npm install` in your CI pipeline is a runtime dependency on that system. When npm is slow, your builds are slow. When npm is compromised, your builds are compromised. When npm decides to pull a package, your builds break. A self-hosted package registry on European infrastructure is a sovereignty boundary. It decouples your build pipeline from a US-operated centralised system. This is not a theoretical concern — it is the concrete difference between being able to freeze your registry at 3 AM when a compromise is reported and waiting for a US-based company to complete their incident response process on US business hours. The same logic applies to your vulnerability databases. osv-scanner can be configured with a local mirror of the OSV database. Trivy can run with a local vulnerability database. Your scanning infrastructure does not need to phone home to US-operated services during every build. ## What this requires Honesty about effort: this is not free. Building the supply chain hardening layer — proxying registry, CI pipeline defences, scanning toolchain, SBOM generation, and Backstage integration for impact assessment — takes 6 to 11 weeks of platform engineering time for an organisation starting from a standard GitLab and Kubernetes setup. Less if GitLab's package registry and Backstage are already deployed. The investment is roughly EUR 43,000 to 80,000 in engineering effort, depending on the complexity of your existing CI pipeline and the number of language ecosystems you need to cover. The return is not measured in cost savings. It is measured in incident response time. When the next axios happens — and it will — the question is whether your platform team can contain it in minutes or whether you spend the next 48 hours in a war room trying to figure out which services are affected. ## Infrastructure versus tooling The axios incident is a useful litmus test. An Internal Developer Platform that includes supply chain defence — a proxying registry, CI-enforced protections, automated scanning, and a service catalogue that can answer "what is affected" — is infrastructure. It absorbs the incident. It gives your team options. It contains the blast radius. A collection of CI/CD scripts that happen to run some security checks is tooling. It might catch the problem. It might not. When it doesn't, you discover exactly how much of your supply chain defence was held together by convention rather than enforcement. The 135 organisations that contacted the attacker's server during the axios window did not lack security awareness. They lacked a control point between the public registry and their build systems. That control point is not a product you buy. It is a platform capability you build — and it is the part of your IDP that earns its keep on the worst days. --- *This article is published on [Clouds of Europe](https://clouds-of-europe.eu), a practitioner community building European cloud independence. We assess infrastructure patterns based on real production experience — no vendor sponsorship, no product agenda.* --- ## How Backstage Helps You Pass a NIS2 Audit URL: https://clouds-of-europe.eu/content/practice/implementation-patterns/how-backstage-helps-you-pass-a-nis2-audit Author: Jurg van Vliet Published: 2026-03-30 Category: Practice Type: Implementation Patterns We published our [NIS2 for Platform Teams](https://clouds-of-europe.eu/nis2-for-platform-teams) guide last year. The most common follow-up question from CTOs and platform leads has been practical: "We understand the requirements. What tooling actually helps us meet them?" The answer is not a single product. NIS2 compliance is a combination of processes, documentation, tooling, and operational discipline. But one component shows up repeatedly as the connective tissue that makes audit responses fast and verifiable: Backstage. Not because Backstage is a security tool. It is not. But because the questions auditors ask — who owns this, what does it depend on, who can access it — are fundamentally catalogue questions. And a well-maintained service catalogue turns hours of scrambling into seconds of querying. This article maps specific NIS2 audit requirements to specific Backstage capabilities. Where Backstage helps, we say so. Where it does not, we say that too. ## What auditors actually ask If you have not been through a NIS2 audit yet, it helps to understand what happens in practice. Auditors do not walk in with a checklist of abstract policy requirements. They ask your platform team to demonstrate things — live, on screen, with real data. Here are the questions that trip up the most organisations, drawn from our experience with European platform teams preparing for and going through NIS2 assessments: **Access control:** - "Show me who has production cluster admin access right now." Not who should have access according to your policy document. Who actually does, today. - "Show me the last time someone's access was reviewed and revoked." - "Walk me through what happens when an engineer leaves. How quickly is access removed across all systems?" **Change management:** - "Show me the deployment pipeline for your production environment." - "Pick a recent production change. Can you trace it from commit to deployment, including who approved it?" **Supply chain:** - "Show me your service dependency register." - "For this service, what external providers does it depend on? Where is the data processed? Under which jurisdiction?" **Business continuity:** - "What is your recovery time objective for your core platform?" - "What happens if your primary cloud region goes down?" The common thread: auditors want to see living systems, not static documents. A PDF that someone wrote six months ago and has not updated since is not a convincing answer. A queryable catalogue that shows current state — and that teams actively maintain as part of their workflow — is. ## Backstage as your service dependency register NIS2 Article 21(2)(d) requires "supply chain security, including security-related aspects concerning the relationships between each entity and its direct suppliers or service providers." In practice, this means you need a documented inventory of every service you run, every external dependency, and the relationships between them. Most organisations attempt this with a spreadsheet. Someone on the platform team creates a Google Sheet with columns for service name, owner, dependencies, and provider jurisdiction. It is accurate for about three weeks. Then a team ships a new service without updating the sheet, someone changes a dependency without telling anyone, and the register drifts from reality. By the time the auditor arrives, the spreadsheet describes an infrastructure that no longer exists. Backstage solves this structurally, not through discipline. The software catalogue in Backstage is a registry of every component, service, API, and resource in your organisation. Each entity is defined in a YAML descriptor file (`catalog-info.yaml`) that lives alongside the source code in the service's repository. The team that owns the service owns the descriptor. When they add a dependency, they update the descriptor. When they change the API, they update the descriptor. The catalogue stays current because it is maintained by the people who know the truth — the teams doing the work. A service entry in the catalogue can include: - **Owner:** Which team is responsible. Not a name in a spreadsheet — a team entity in Backstage that links to real people. - **Dependencies:** Which other services this one consumes. Expressed as relations between catalogue entities, so you can traverse the graph. - **APIs:** What this service exposes and what it consumes. Defined as API entities with specifications (OpenAPI, gRPC, AsyncAPI). - **External dependencies:** Cloud provider services, SaaS integrations, third-party APIs. Each can be modelled as a Resource entity with metadata about the provider and jurisdiction. - **Data classification:** Custom metadata fields indicating what kind of data this service processes and where. - **Provider jurisdiction:** Custom annotations documenting where the service and its dependencies are hosted, and under which legal framework. When an auditor asks "show me your service dependency register," you open Backstage. You show them a live graph of services, their owners, and their dependencies. You click into any service and show them its external dependencies, data classification, and provider jurisdiction. You demonstrate that this information is version-controlled — every change to a `catalog-info.yaml` is a Git commit with an author, a timestamp, and a review trail. This is qualitatively different from a spreadsheet. It is queryable. It is version-controlled. It is maintained by the teams closest to the truth. And it is auditable by design. ### What the YAML looks like in practice This is not abstract. Here is what a `catalog-info.yaml` looks like for a service with NIS2-relevant metadata: ```yaml apiVersion: backstage.io/v1alpha1 kind: Component metadata: name: payment-service description: Processes customer payments annotations: backstage.io/techdocs-ref: dir:. nis2/data-classification: confidential nis2/provider-jurisdiction: EU (Frankfurt) nis2/recovery-time-objective: 4h nis2/last-risk-assessment: "2026-02-15" spec: type: service lifecycle: production owner: team-payments providesApis: - payment-api consumesApis: - customer-api - billing-api dependsOn: - resource:default/postgresql-eu-west - resource:default/stripe-payment-gateway ``` The `nis2/` annotations are custom. Backstage lets you define whatever metadata your organisation needs. The point is that this file lives in the same repository as the service code, is reviewed in the same pull request workflow, and is automatically ingested into the catalogue. ## Backstage for ownership and access documentation Every entity in Backstage has an owner. This is not optional — the catalogue schema requires it. When an auditor asks "who owns this service?" you answer in seconds, not hours. This matters more than it might seem. In organisations with 50+ services, the question "who owns X?" can take surprisingly long to answer if the information lives in people's heads, Confluence pages, or a spreadsheet that predates three reorganisations. We have seen platform teams spend 30 minutes in an audit tracking down the owner of a service that turned out to be maintained by a team that was renamed twice. Backstage solves this by making ownership a structural property of the catalogue. Every Component, API, and Resource has an owner. The owner is itself an entity — a Group or a User — so you can click through to see who is on that team and what else they own. For access control documentation, Backstage is not the system of record — your identity provider is. But Backstage can surface access information in context. Plugins exist to integrate with identity providers (Okta, Azure AD, Keycloak) and display who has access to what, alongside the service they are looking at. The Backstage RBAC plugin (available in both the community and commercial distributions) lets you visualise and manage permissions within the portal itself. More importantly for NIS2, it provides an auditable record of who can see and modify what within the developer portal — which is itself a system that auditors may ask about. The practical workflow during an audit: 1. Auditor asks: "Who owns the payment service?" 2. You open Backstage, search for `payment-service`, show the owner: `team-payments`. 3. Click through to `team-payments`, see the members, their roles. 4. From the same page, show the other services this team owns. 5. Show the last changes to the service's catalogue entry — all in Git history. Total time: under 60 seconds. Compare that to an organisation without a catalogue, where the answer involves checking three Confluence spaces, asking on Slack, and hoping someone remembers. ## Backstage for SBOM and dependency scanning NIS2's supply chain requirements extend beyond infrastructure dependencies to software dependencies. Article 21(2)(d) combined with ENISA's implementing guidance means that auditors increasingly expect Software Bill of Materials (SBOM) documentation and evidence of dependency vulnerability management. Backstage does not generate SBOMs. But it serves as the integration point where SBOM data becomes accessible and actionable. **TechDocs for compliance documentation.** Backstage's TechDocs feature renders Markdown documentation stored alongside your code directly in the portal. For NIS2, this means you can maintain compliance-relevant documentation — risk assessments, data processing records, recovery procedures — as code, reviewed through pull requests, and accessible through the same interface where teams find everything else about a service. Documentation that lives where developers already work actually gets read and updated. Documentation in a separate compliance tool does not. **Security plugin ecosystem.** The Backstage plugin ecosystem includes integrations with dependency scanning tools that your platform team likely already runs: - **Snyk plugin:** Surfaces vulnerability findings per service directly in the catalogue. - **Dependabot/Renovate integration:** Shows dependency update status alongside service metadata. - **Security scorecard plugins:** Aggregate security posture metrics (dependency freshness, known vulnerabilities, scan coverage) per service and per team. - **SBOM plugins:** Display generated SBOMs (from tools like Syft or Trivy) within the service's Backstage page. The pattern here is consistent: Backstage does not replace your security tooling. It aggregates the outputs and makes them discoverable in context. When an auditor asks about vulnerability management for a specific service, you can show scanning results, dependency status, and SBOM data from the same page that shows ownership and dependencies. This matters for audit efficiency. Auditors do not want to log into five different tools to assess one service. A single interface that shows ownership, dependencies, security posture, and documentation for each service makes the audit faster and more thorough — which works in your favour. ## What Backstage does not solve This is the section that matters most, because adopting Backstage and believing your NIS2 compliance is handled would be a dangerous mistake. **Backstage is a catalogue, not a security tool.** It documents who has access. It does not enforce access controls. It shows dependencies. It does not scan them for vulnerabilities. It surfaces security findings from other tools. It does not generate those findings. The distinction is critical. **Backstage does not detect incidents.** When something goes wrong at 3 AM, Backstage provides context — who owns the affected service, what it depends on, who to call. That is valuable during incident response. But it is not your monitoring, alerting, or SIEM system. You still need Prometheus for metrics, a log aggregation stack, and a security event pipeline. Backstage tells you what you are looking at. Your observability stack tells you something is wrong. **Backstage does not enforce change management.** It can display deployment history and link to CI/CD pipelines, but the actual controls — branch protection, required reviews, deployment gates — live in your Git platform and your CD tooling (Flux, ArgoCD). Backstage shows you the audit trail. Your CI/CD pipeline creates it. **Backstage does not prove business continuity.** You can document RTOs and recovery procedures in TechDocs. You can annotate services with their recovery classifications. But Backstage cannot test your disaster recovery, validate your backups, or failover your services. Those are operational capabilities that require separate investment. **Backstage requires sustained investment.** As we detailed in our [Building an IDP from Open Source](https://clouds-of-europe.eu/building-an-idp-from-open-source) article, Backstage is a framework, not a product. Getting it to a state where it actually helps during an audit — catalogue populated, plugins configured, documentation templates in place, teams actively maintaining their entries — takes 4-8 weeks of dedicated engineering time for initial setup, and 0.25-0.5 FTE for ongoing maintenance. If you deploy Backstage, populate the catalogue once, and then let it drift for six months, you will be worse off than having no catalogue at all. An auditor who sees a stale Backstage instance will rightly question whether any of your documentation is current. A catalogue that shows "last updated 8 months ago" on half its entries actively undermines your compliance posture. To be direct: Backstage is one component in a NIS2 compliance toolkit. A valuable one, but not the whole answer. You still need: - **Observability** (Grafana stack, Prometheus, Loki) for detection and investigation - **SIEM or security event pipeline** for incident detection and correlation - **Identity provider** (Keycloak, Okta, Azure AD) for access control enforcement - **GitOps tooling** (Flux, ArgoCD) for auditable change management - **Backup and DR tooling** for business continuity - **Vulnerability scanning** (Trivy, Snyk, Grype) for supply chain security Backstage ties these together. It is the index, not the library. ## Practical setup: a minimal NIS2-ready Backstage configuration If you are setting up Backstage with NIS2 compliance as an explicit goal, here is what a minimal viable configuration looks like. This is not the full developer portal experience — it is the subset that directly supports audit readiness. ### 1. Service catalogue with NIS2 metadata Define a catalogue entity schema that includes the metadata auditors will ask about: - **Owner** (required by default) - **Lifecycle stage** (production, staging, deprecated) - **Data classification** (custom annotation) - **Provider jurisdiction** (custom annotation) - **Recovery time objective** (custom annotation) - **External dependencies** modelled as Resource entities with provider and jurisdiction metadata Establish a convention that every production service has a `catalog-info.yaml` with these fields populated. Enforce it through CI — a pipeline check that fails if a production service lacks required NIS2 annotations. Automation beats discipline every time. ### 2. TechDocs with compliance templates Create TechDocs templates for: - **Service risk assessment:** What data does this service process? What are the threats? What mitigations are in place? - **Recovery procedure:** How do you restore this service from backup? What is the expected recovery time? When was this last tested? - **Incident response context:** During an incident involving this service, what are the first diagnostic steps? Who should be contacted? Store these as Markdown files alongside the service code. Render them in Backstage via TechDocs. Teams fill in the templates as part of their service onboarding process. ### 3. Identity provider integration Connect Backstage authentication to your organisation's identity provider. This accomplishes two things: - Users in Backstage correspond to real identities in your directory, which means ownership data in the catalogue maps to real people. - Backstage itself has an auditable authentication trail — who accessed the portal, when, and what they viewed. ### 4. API catalogue for external dependencies Model external service dependencies as first-class entities in the catalogue. Every SaaS integration, every third-party API, every cloud-managed service should be a Resource entity with: - Provider name - Jurisdiction (headquarters and data processing location) - Data access level (what can this provider see?) - Contract expiry date - Link to data processing agreement This is the supply chain register that auditors will ask for. Building it in the Backstage catalogue — rather than in a standalone spreadsheet — means it is linked to the services that depend on it. When an auditor picks a service and asks "what third-party providers does this depend on?", you can traverse the dependency graph and show them. ### 5. Security plugin integration At minimum, install and configure: - A vulnerability scanning plugin that shows findings per service - A dependency freshness tracker - An SBOM display plugin if your build pipeline generates SBOMs These do not need to be sophisticated on day one. The point is that security posture data is visible in the same place as ownership and dependency data, so auditors can assess a service holistically from a single interface. ## The audit scenario Here is what a good NIS2 audit interaction looks like when you have Backstage set up properly: **Auditor:** "Show me your service dependency register." You open Backstage. You show the full catalogue — 47 services, 12 APIs, 23 external resources. You filter by lifecycle: production. 31 services. You click one. You show its dependencies, its owner, its data classification, its provider jurisdiction. You show that the catalogue entry was last updated three days ago as part of a pull request that added a new dependency. **Auditor:** "Who owns this service, and who has production access?" You click the owner: `team-platform`. You show the team members. You click through to the identity provider integration showing current access grants. You demonstrate that access was last reviewed two weeks ago. **Auditor:** "Show me the deployment pipeline and trace a recent change." You click the CI/CD plugin tab. You show the last ten deployments, each linked to a Git commit. You click one, it opens the merge request in GitLab — author, reviewer, approval, pipeline run, deployment timestamp. Full traceability. **Auditor:** "What is the recovery time objective for this service?" You point to the annotation: `nis2/recovery-time-objective: 4h`. You click through to the TechDocs recovery procedure, which documents the restore process and shows it was last tested six weeks ago. Total time for this exchange: ten minutes. Without a catalogue, the same questions take an hour or more, involve multiple people, and produce answers the auditor trusts less. ## Start with the catalogue If you are preparing for a NIS2 audit and considering Backstage, here is our recommendation: start with the catalogue and TechDocs. Do not try to deploy every plugin and integration on day one. **Week 1-2:** Deploy Backstage with authentication. Define your `catalog-info.yaml` schema with NIS2 annotations. Write the CI check that validates required fields. Start populating — begin with your ten most critical production services. **Week 3-4:** Add TechDocs templates for risk assessment and recovery procedures. Get the first three teams to fill them in. Iterate on the templates based on their feedback. **Month 2:** Model external dependencies as Resource entities. This is your supply chain register. Add the identity provider integration. **Month 3:** Add security plugins. Connect vulnerability scanning data. Run a mock audit internally — have someone play the auditor and go through the questions listed above. Find the gaps. This is not fast. It should not be fast. A catalogue that is rushed into existence and populated with garbage data is worse than no catalogue at all. The value comes from accuracy and currency, not from speed of deployment. Backstage is not the answer to NIS2 compliance. But it is the connective tissue that makes the rest of your compliance tooling auditable, discoverable, and fast to query. When the auditor is sitting across from your platform team, that speed and confidence is what separates a clean audit from a finding. --- *This article is published on [Clouds of Europe](https://clouds-of-europe.eu), a practitioner community building European cloud independence. We assess infrastructure tooling based on real production experience — no vendor sponsorship, no product agenda.* --- ## NIS2 for Platform Teams: What Changes in Practice URL: https://clouds-of-europe.eu/content/strategy-transition/strategic-planning/nis2-for-platform-teams-what-changes-in-practice Author: Jurg van Vliet Published: 2026-03-30 Category: Strategy & Transition Type: Strategic Planning The NIS2 Directive (Directive (EU) 2022/2555) entered enforcement across EU member states in October 2024, and national transpositions are landing now. If you run infrastructure for a European organisation of any significant size, this affects you — probably more than your legal team has communicated. The problem with most NIS2 guidance is that it stops at the compliance-officer level. It talks about "appropriate technical measures" and "supply chain risk management" without ever getting specific about what that means for the team running Kubernetes clusters, managing cloud accounts, and operating CI/CD pipelines at 2am. This article maps NIS2 requirements to the infrastructure decisions your platform team makes every day. No policy abstractions. Specific technical implications. ## You are probably an "essential" or "important" entity NIS2 classifies organisations into two categories: **essential entities** and **important entities**. The penalties and oversight differ significantly, but both categories face real obligations. Here is where most organisations underestimate their exposure: **Essential entities** include energy, transport, banking, health, drinking water, digital infrastructure, ICT service management (B2B), and public administration. If you are a managed service provider, a cloud service provider, a data centre operator, or a DNS provider — you are an essential entity. Full stop. **Important entities** cover a broader set: postal services, waste management, chemicals, food, manufacturing, digital providers, and research. The "digital providers" category is where it gets interesting. Online marketplaces, search engines, and social networking platforms are explicitly included, but so are many SaaS businesses that member states interpret broadly. The threshold: **medium-sized or larger** (50+ employees or EUR 10M+ turnover). If you are reading this article, your organisation almost certainly qualifies. **Why this matters for platform teams:** The level of scrutiny your infrastructure will face depends on this classification. Essential entities face proactive supervision — regulators can audit you without a triggering incident. Important entities face reactive supervision — audits happen after an incident or report. Either way, your platform team needs to be ready to answer detailed questions about how your infrastructure operates. Most CTOs we speak with assumed NIS2 was primarily about their security team. It is not. The directive's technical requirements land squarely on whoever controls the infrastructure, the deployment pipeline, and the observability stack. That is your platform team. ## Supply chain documentation: your cloud provider is your supply chain Article 21(2)(d) of NIS2 requires "supply chain security, including security-related aspects concerning the relationships between each entity and its direct suppliers or service providers." Read that again. Your cloud provider is a direct supplier. Your managed database service is a direct supplier. Your CDN, your DNS provider, your container registry, your secrets manager — all direct suppliers. **What this means in practice:** You need a documented inventory of every third-party service your infrastructure depends on. Not a spreadsheet that someone made once and forgot about. A living inventory that answers these questions: - **Jurisdiction:** Where is this provider headquartered? Where is the data processed? Under which legal frameworks can the provider be compelled to disclose data? - **Substitutability:** If this provider becomes unavailable (or legally problematic), what is the migration path? How long would it take? - **Access scope:** What access does this provider have to your systems and data? Can they read your data at rest? In transit? - **Incident notification:** How does this provider notify you of security incidents? What are their contractual SLAs for notification? For organisations running on AWS, Azure, or Google Cloud, the jurisdiction question is unavoidable. These are US-headquartered companies subject to the CLOUD Act (Clarifying Lawful Overseas Use of Data Act, 2018), which allows US authorities to compel disclosure of data stored abroad. NIS2 does not explicitly prohibit using US providers, but it does require you to document this risk and demonstrate that you have assessed it. **The practical output your platform team needs to produce:** 1. **Service dependency register** — Every external service, its function, its jurisdiction, its data access level. This is not a one-time exercise. Automate it. If you run Kubernetes, your cluster's dependencies are a good starting point: container runtime, CNI plugin, CSI drivers, ingress controllers, certificate management, DNS. Many of these phone home or depend on external services. 2. **Risk assessment per supplier** — Document the impact if each supplier is compromised, becomes unavailable, or is compelled to act against your interests. Be specific. "AWS eu-west-1 becomes unavailable for 72 hours" is a scenario. "Our container images are in ECR and we have no secondary registry" is a finding. 3. **Contractual review** — Does your agreement with each provider include incident notification clauses? Data processing agreements? Do they meet NIS2 requirements for notification timelines? Most hyperscaler standard agreements do not. This is the kind of documentation that auditors will ask for first because it is concrete and verifiable. You either have a current service dependency register or you do not. ## Incident reporting: 24 hours is not much time NIS2 introduces strict incident reporting timelines that your platform team needs to be able to meet — not in theory, but with actual tooling and processes. **The timeline:** - **24 hours** — Early warning to the competent authority (national CSIRT or relevant body). This notification must include whether the incident is suspected to be caused by unlawful or malicious acts, and whether it could have cross-border impact. - **72 hours** — Full incident notification. This must include an initial assessment of the incident's severity and impact, plus indicators of compromise where available. - **1 month** — Final report with detailed description, root cause analysis, mitigation measures applied, and cross-border impact if applicable. **What your observability stack needs to support:** The 24-hour early warning is the critical constraint. That means from the moment a significant incident begins, your team has 24 hours to detect it, assess it, and file a structured notification with the relevant authority. Work backwards from that requirement: **Detection latency** — How long does it take between a security event occurring and your team knowing about it? If your log aggregation has a 15-minute delay, and your alerting checks every 30 minutes, and your on-call engineer takes 30 minutes to respond, you have already spent more than an hour before anyone starts investigating. For a 24-hour reporting deadline, that margin matters. **Log completeness** — The 72-hour notification requires indicators of compromise and initial root cause. You cannot produce these if your logging is incomplete. At minimum, your platform needs: - API server audit logs (every authentication and authorisation decision) - Network flow logs (or at least ingress/egress at the cluster boundary) - Container runtime events (image pulls, exec into containers, privilege escalations) - Access logs for every externally-facing service - Change logs for infrastructure-as-code (who deployed what, when, from which commit) **Structured incident classification** — NIS2 defines a "significant incident" as one that has caused or is capable of causing severe operational disruption or financial loss, or has affected or is capable of affecting other natural or legal persons by causing considerable damage. Your incident severity framework needs to map to this definition. Severity 1 in your runbook should correspond to "significant incident" under NIS2, and the escalation path should include regulatory notification as a step. **Practical recommendation:** Build a NIS2 incident reporting template into your incident response runbook. When your on-call engineer declares a significant incident, the template should auto-populate with: - Timestamp of detection - Services affected (from your service catalogue) - Data classifications involved (from your data inventory) - Initial indicators of compromise (from your SIEM or log analysis) - Whether cross-border impact is possible (based on where your users and data reside) The engineer should not have to think about NIS2 compliance during an incident. The process should produce the required outputs by default. ## What auditors will actually ask your platform team Legal teams tend to prepare for NIS2 audits by assembling policy documents: information security policies, risk management frameworks, business continuity plans. These matter, but they are table stakes. The auditor already assumes you have them. Where audits get interesting — and where organisations fail — is when the auditor moves from policy to implementation. Here is what your platform team should expect: **Access control:** - "Show me who has production cluster admin access right now." (Not who is supposed to — who actually does.) - "Show me the last time someone's access was reviewed and revoked." - "How do you handle access for third-party contractors?" - "Walk me through what happens when an engineer leaves the company. How quickly is access revoked across all systems?" If your answer to any of these requires someone to manually check multiple systems, you have a finding. Auditors expect centralised identity management with automated provisioning and deprovisioning. If your Kubernetes RBAC is managed through static YAML files that someone applies manually, that is a gap. **Change management:** - "Show me the deployment pipeline for your production environment. Who can push changes, and what gates exist?" - "Show me a recent production change. Can you trace it from code commit to deployment, including who approved it?" - "How do you handle emergency changes that bypass the normal process?" GitOps workflows score well here. If every production change is a pull request in a Git repository, reviewed by a second person, and applied by Flux or ArgoCD, you have a verifiable audit trail by default. If you are still doing `kubectl apply` from laptops, expect follow-up questions. **Logging and monitoring:** - "Show me your log retention policy. How far back can you investigate an incident?" - "Show me how you would detect unauthorised access to your data stores." - "What is your mean time to detect a security event?" The retention question catches many organisations. NIS2 does not specify a retention period, but auditors will ask whether your retention aligns with the 1-month final report timeline. If you only retain logs for 14 days, you cannot produce a root cause analysis for an incident that started 20 days ago. **Business continuity:** - "What is your recovery time objective for your core platform?" - "When was the last time you tested your disaster recovery procedure?" - "If your primary cloud region became unavailable, what happens?" The last question is where supply chain documentation and business continuity intersect. If your entire stack runs in a single cloud provider's single region, your honest answer might be "we would be offline for an extended period." NIS2 does not require multi-cloud, but it does require you to have assessed this risk and have a documented response plan. ## The practical checklist Here is what your platform team should have in place, mapped to NIS2 requirements. This is not exhaustive — it covers the areas where platform teams are most likely to be directly involved. ### Logging and detection - [ ] Centralised log aggregation with defined retention period (minimum 30 days, 90 days recommended) - [ ] API server audit logging enabled with request and response metadata - [ ] Network flow logs at cluster boundary - [ ] Container runtime security events (image pulls, exec, privilege escalation) - [ ] Alerting on authentication failures, privilege escalations, and unusual access patterns - [ ] Documented mean time to detect (MTTD) for security events - [ ] Log integrity protection (immutable storage or cryptographic chaining) ### Access control - [ ] Centralised identity provider integrated with all infrastructure components - [ ] Role-based access control with documented role definitions - [ ] Just-in-time or time-limited access for production environments - [ ] Automated deprovisioning when team members leave or change roles - [ ] Regular access reviews (quarterly at minimum) with documented outcomes - [ ] Multi-factor authentication for all infrastructure access - [ ] Separate credentials for CI/CD pipelines (no shared service accounts with human users) ### Supply chain inventory - [ ] Complete register of external services and providers - [ ] Jurisdiction documentation for each provider (headquarters, data processing locations) - [ ] Risk assessment for each provider, including substitutability analysis - [ ] Contractual review for incident notification and data processing terms - [ ] Automated dependency scanning for software supply chain (container images, libraries) - [ ] SBOM (Software Bill of Materials) generation for deployed services - [ ] Regular review cadence (quarterly) for supply chain register updates ### Incident response - [ ] Documented incident response runbook with NIS2 reporting steps integrated - [ ] Pre-built notification templates aligned with 24h/72h/1month timelines - [ ] Defined incident severity levels mapped to NIS2 "significant incident" criteria - [ ] On-call rotation with documented escalation paths including regulatory notification - [ ] Tested incident response process (tabletop exercise within the last 12 months) - [ ] Designated contact for national CSIRT/competent authority communication - [ ] Post-incident review process that produces the information required for the 1-month final report ### Business continuity - [ ] Documented recovery time objectives (RTO) and recovery point objectives (RPO) - [ ] Tested backup and restore procedures for stateful services - [ ] Disaster recovery plan that addresses single-provider and single-region failure - [ ] Documented dependencies on third-party services for recovery - [ ] Annual DR test with documented results and remediation actions ## What to do next If you are starting from scratch, prioritise in this order: **Week 1-2: Supply chain inventory.** Build the service dependency register. This is the most visible gap and the first thing auditors request. Start with your cloud provider accounts, then work through DNS, CDN, container registry, secrets management, monitoring, and CI/CD tooling. **Week 3-4: Incident response integration.** Take your existing incident response runbook and add NIS2 reporting steps. Build the notification templates. Identify your national competent authority and CSIRT. Run a tabletop exercise with the extended timeline requirements. **Month 2: Logging and access control gaps.** Audit your current logging against the checklist above. Enable what is missing. Review access control — can you answer the auditor questions listed above with current tooling, or do you need to consolidate? **Month 3: Business continuity testing.** If you have not tested your disaster recovery procedure in the last 12 months, schedule it. Document the results. Identify gaps and build a remediation plan. None of this requires buying new products. It requires your platform team to document what they have, identify gaps, and address them systematically. The organisations that handle NIS2 well are the ones where the platform team owns this work directly rather than waiting for legal or compliance to translate requirements into technical tasks. NIS2 is not a security-team-only problem. It is an infrastructure problem. Your platform team is best positioned to solve it. --- ## One Maintainer Away from Disaster — How Foundations Build the Safety Net URL: https://clouds-of-europe.eu/content/strategy-transition/collaborative-infrastructure/one-maintainer-away-from-disaster-how-foundations-build-the-safety-net Author: Jurg van Vliet Published: 2026-03-23 Category: Strategy & Transition Type: Collaborative Infrastructure Tags: cncf, kubernetes, linuxfoundation, opensource, sovereignty, sustainability In 2022, every Kubernetes cluster on earth depended on a database maintained by one person. etcd — the distributed key-value store that holds every piece of cluster state, every secret, every deployment manifest — had a single active maintainer. Marek Siarkowicz at Google was it. The others had moved on. Gyuho Lee went to Amazon and stopped participating. Sam Batschelet at Red Hat went quiet. Piotr Tabor at Google and Sahdev P Zala at IBM contributed occasionally, but "occasionally" does not fix critical data inconsistency bugs in a system that underpins the entire cloud-native ecosystem. etcd v3.5 had shipped with multiple data corruption bugs. Not theoretical risks — actual data inconsistency. The kind that silently destroys cluster state. Fixes took months because one person cannot review their own patches, cannot achieve the supermajority required for governance decisions, and — in a detail that should alarm everyone — could not even access the CNCF helpdesk. But here is where the story diverges from the usual open-source tragedy. The situation escalated to the Kubernetes Steering Committee under the subject line "Worrying state of Etcd community." The CNCF Technical Oversight Committee got involved. The resolution was the creation of SIG-etcd in 2023, elevating the project to first-class Kubernetes citizen status with formal governance, dedicated contribution paths, and institutional accountability. New maintainers were recruited. Review capacity was restored. The data corruption bugs were fixed. The foundation caught it. Not fast enough — but it caught it. And the mechanisms it used are worth understanding, because they represent a model for sustaining public technology infrastructure. ## The bus factor is real Software engineers have a grim metric called the "bus factor" — how many people need to disappear before a project dies. For etcd in 2022, it was one. For thousands of critical open-source projects right now, it is one. The CNCF landscape lists over a thousand projects. A project becomes critical infrastructure — adopted by thousands of companies, embedded in production systems that handle billions in transactions — and the people who built it get nothing except more issues filed against them. This is the problem foundations exist to solve. ## What happens without a foundation: the xz backdoor The most dramatic illustration arrived in 2024 with the xz/liblzma backdoor. xz was not a CNCF project. It was not governed by any foundation. It was maintained by one person, Lasse Collin, essentially alone for years. He was visibly burned out. A contributor named "Jia Tan" appeared, built trust over two years through legitimate contributions, and gradually inserted a sophisticated backdoor into the build system — one designed to compromise SSH authentication on every Linux system that linked against liblzma. The attack succeeded precisely because the bus factor was one. There was no second pair of eyes with sufficient context to catch the changes. The social engineering worked because a burned-out maintainer was desperate for help, and there was no institution — no foundation, no governance body, no formal contributor pipeline — to provide it. The backdoor was discovered by accident. Andres Freund at Microsoft noticed SSH connections were 500 milliseconds slower than expected and investigated. Not a code review. Not a security audit. A performance anomaly. Compare this to etcd. Same bus factor. Same risk. But etcd had the CNCF — an institution that could escalate, intervene, restructure, and recruit. xz had nothing. The difference is not the code. The difference is the institution around it. ## Log4Shell: the response that changed everything In December 2021, a critical remote code execution vulnerability was disclosed in Apache Log4j. CVE-2021-44228 — Log4Shell — allowed unauthenticated remote code execution on any system that logged attacker-controlled input, which turned out to be most of the internet. Log4j was maintained by a handful of volunteers. When the vulnerability was disclosed, those volunteers worked around the clock to produce fixes while the entire technology industry — companies with trillion-dollar market capitalisations — waited for patches from people who were not being paid. But Log4Shell did not just expose a vulnerability. It triggered an institutional response. The Linux Foundation launched the Open Source Security Foundation (OpenSSF) mobilisation plan, backed by a $150 million commitment from major technology companies. The Alpha-Omega project was created specifically to fund security work on critical open-source projects. The Scorecard project began systematically evaluating the security posture of open-source dependencies. The crisis was real. But so was the response. And the response was only possible because a foundation — the Linux Foundation — had the institutional weight to convene the industry, allocate resources, and create lasting structures rather than one-time fixes. ## The Linux kernel: what foundation-scale investment looks like The Linux kernel is the proof that this model works at scale. Thousands of paid contributors from hundreds of companies, with the Linux Foundation providing the institutional backbone — governance, infrastructure, legal protection, and coordination. This did not happen by accident. It happened because the Linux Foundation made it possible for companies to invest in shared infrastructure without surrendering competitive advantage. The foundation provides neutral ground. Companies that compete fiercely in the market collaborate on the kernel because the foundation structure makes that collaboration safe and productive. Even Linux is not immune to risk. The kernel.org compromise of 2011 went undetected for seventeen days. The "hypocrite commits" incident in 2021 — when University of Minnesota researchers deliberately submitted flawed patches — exposed how thin the review layer could be. But in both cases, the institutional response was swift: infrastructure rebuilt, policies updated, contributor vetting strengthened. The foundation absorbed the shock and came back stronger. That resilience is not a property of the code. It is a property of the institution. ## ReiserFS: the cost of having no safety net ReiserFS offers the starkest illustration. Hans Reiser created the filesystem, maintained it as part of the Linux kernel, and was convicted of murder in 2008. The filesystem stagnated for sixteen years — a zombie in the codebase — because nobody else had the context or mandate to maintain it. It was finally removed from the kernel in 2024. "Maintainer convicted of murder" is extreme. But it sits on a spectrum with "maintainer changes jobs," "maintainer has a child," "maintainer gets burned out." All normal human events. All capable of killing a project that millions depend on — unless an institution exists to ensure continuity. ## What CNCF and Linux Foundation governance actually provides Foundations are not charities. They are infrastructure for infrastructure. Here is what the CNCF provides, concretely: **Graduation criteria that enforce maintainer depth.** A CNCF project cannot graduate from sandbox to incubating to graduated without demonstrating healthy contributor diversity, documented governance, and a security audit. These are not suggestions. They are gates. etcd's crisis happened partly because these mechanisms had not yet been applied retroactively to projects that predated them. That gap has been closed. **Technical Oversight Committee (TOC) intervention.** The TOC is not decorative. When etcd was failing, the TOC acted — restructuring the project's governance, creating SIG-etcd, and ensuring institutional support. This is the immune system in action: detecting a failing project and mobilising resources before it collapses. **Security audits and vulnerability response.** CNCF funds third-party security audits for graduated projects. The Linux Foundation's OpenSSF provides tooling, funding, and coordination for security across the broader ecosystem. These are not things individual maintainers can do alone. **Neutral governance for competing contributors.** When Google, Red Hat, and VMware all contribute to Kubernetes, the CNCF provides the legal and organisational framework that makes that possible. Without it, every contribution becomes a negotiation between corporate legal departments. **Contributor pipelines.** Programs like LFX Mentorship and Google Summer of Code, coordinated through foundation infrastructure, create pathways for new maintainers. This is how you solve the bus factor — not by hoping someone shows up, but by building the pipeline that ensures they do. ## Contributing back: from TAG Environmental Sustainability to KEIT Foundations do not just protect projects. They create the spaces where new ideas take root. Our CTO, Flavia, served as a SIG Tech Lead for CNCF's TAG Environmental Sustainability — the group working on carbon-aware computing, sustainability metrics, and environmental impact measurement for cloud-native infrastructure. That work — the conversations, the specifications, the cross-company collaboration that only a foundation can convene — directly inspired KEIT, the Kubernetes Emissions Insights Tool we built and open-sourced. KEIT estimates the carbon emissions of a Kubernetes cluster by combining Kepler for energy measurement, ElectricityMaps for grid carbon intensity, and Boavizta for hardware embodied emissions. It exists because a foundation created the space to think about the problem, connected the people working on it, and established the shared vocabulary that made a practical tool possible. This is the foundation model working as designed: participate, learn, contribute back. The investment flows in both directions. ## Open source is public technology Roads, bridges, water systems — we call these public infrastructure and fund them accordingly. Open-source software that runs the global economy deserves the same treatment. The CNCF and Linux Foundation are, in effect, the governance bodies for public technology. They maintain the roads that every company drives on. When those roads are well-funded — when companies invest in foundation membership, employ maintainers, participate in governance, contribute upstream — the entire ecosystem benefits. When they are neglected, we get xz. Three things every European organisation running open-source infrastructure should do: **Fund the foundations that govern your dependencies.** Not as a marketing exercise. Not for the conference badge. Because your production infrastructure depends on software that needs institutional support to survive. CNCF membership, Linux Foundation membership, OpenSSF participation — these are investments in the reliability of your own systems. **Contribute people, not just money.** Foundations need participants. SIGs and TAGs need leads. Working groups need reviewers. The most valuable contribution is sustained human engagement — engineers who show up, review code, mentor newcomers, and build the maintainer depth that prevents the next etcd crisis. **Map your dependencies and their governance.** Know which of your critical dependencies are governed by a foundation and which are one maintainer away from disaster. For the governed ones, invest in them. For the ungoverned ones, help bring them into a foundation — or build your own contingency plan. The threat to open-source infrastructure is not technical. It is institutional. The good news is that the institutions exist, they work, and they are getting stronger. The question is whether we invest in them before the next crisis — or only after. --- #kubernetes #cncf #linuxfoundation #opensource #sovereignty #sustainability --- ## A Sovereign Central Intelligence URL: https://clouds-of-europe.eu/content/policy-sovereignty/provider-ecosystem/a-sovereign-central-intelligence Author: Jurg van Vliet Published: 2026-03-19 Category: Policy & Sovereignty Type: Provider Ecosystem We are building an AI-powered commercial intelligence platform. It extracts structured knowledge from our content, scores it for quality using multiple independent models, and feeds a training system that helps our sales team articulate our value proposition. The core of this pipeline runs on European infrastructure with open-weight models. But we also use Claude Code and Anthropic models — deliberately, in roles where they add value without creating dependency. Here is how we designed it, and why. ## The problem We have a growing library of content: marketing documents, published articles, case studies, competitive positioning. This content contains commercial insights — perspective shifts that help prospects see their own situation differently. Manually curating these insights does not scale. As we publish more, the gap between what exists and what our sales team actually uses widens. We needed a system that automatically extracts, validates, and maintains a structured knowledge base from our content sources. The catch: this system touches our most sensitive commercial intelligence. It must run on infrastructure we control, using models we can replace. ## The architecture Our pipeline has five stages, each with a deliberate model choice. ### Stage 1: Extraction A large language model reads each content source and extracts structured entities — insights with reframes, evidence, stakeholder tailoring, and triggers. This is the most demanding task: it requires deep comprehension of what makes a commercial insight useful versus what is merely a summary. **Model:** Qwen 3.5 397B-A17B, running on Scaleway Generative APIs in Paris. Open-weight (Apache 2.0), Mixture-of-Experts architecture (only 17 billion parameters active per token, making it cost-efficient despite 397 billion total). Tier S on the Onyx self-hosted LLM leaderboard. **Why this model:** It has the best structured output quality of any open-weight model available to us. The MoE architecture keeps costs reasonable for processing large document sets. ### Stage 2: Field validation Code, not a model. Strict schema validation rejects malformed records — unknown fields, missing required fields, invalid categories. This catches the inevitable quirks of LLM output (misspelled field names, objects where strings are expected) before they enter the database. **Why code, not a model:** Validation rules are deterministic. There is no reason to spend inference tokens on something a few lines of Python handles perfectly. ### Stage 3: Confidence scoring Three architecturally different models independently score each extracted insight on a 0.0-1.0 scale. The key principle: the scorers must be different from the extractor. Same model scoring its own output is biased. Different architectures trained on different data with different biases give genuinely independent assessments. **Models:** 1. DeepSeek R1-Distill Llama 70B — reasoning specialist, evaluates through chain-of-thought 2. Llama 3.3 70B Instruct — best instruction-following in its class, reliably applies the scoring rubric 3. Gemma 3 27B — Google architecture, cheapest scorer, catches what the other two miss All three run on Scaleway Generative APIs. Open-weight. European infrastructure. **Decision logic:** Take the median score (resists one outlier). If all three agree above 0.7, auto-accept. If all three agree below 0.4, auto-reject. If they disagree significantly (spread above 0.3), flag for review. ### Stage 4: Triage coordination Flagged items — the ones where the three scorers disagreed — need a more nuanced evaluation. Is this a genuine insight that one scorer undervalued, or is it borderline content that happened to get one generous score? **Model:** Claude Haiku 4.5 (Anthropic). Fast, cost-effective, handles classification at 90% of the quality of larger models. Processes the flagged queue in bulk. ### Stage 5: Arbitration The truly ambiguous cases — where even Haiku cannot make a confident call — escalate to a final arbiter. **Model:** Claude Sonnet 4.6 (Anthropic). The deepest reasoning available. Only touches the small subset of cases that survived four previous stages of filtering. ## The sovereignty posture The core pipeline — extraction, validation, scoring, accept/reject — runs entirely on European infrastructure using open-weight models. No data leaves the European jurisdiction for these functions. Anthropic models handle triage and arbitration. These are quality-enhancement functions, not critical-path functions. If Anthropic becomes unavailable tomorrow: - The extraction continues (Qwen on Scaleway) - The scoring continues (DeepSeek, Llama, Gemma on Scaleway) - Auto-accept and auto-reject continue (code logic, no model needed) - Flagged items queue for human review instead of Haiku triage - Ambiguous items go directly to human review instead of Sonnet arbitration The pipeline degrades gracefully. More human review, same data quality. No data loss, no service interruption, no architectural change required. ## The model selection policy We formalized three rules: 1. **No OpenAI models.** Not for extraction, scoring, embedding, or any other function. 2. **Open source first.** Prefer open-weight models that can be self-hosted if needed. 3. **European second.** When choosing between equivalent open-source models, prefer European-origin models. Anthropic models are used where they add unique value — and only in roles where they can be replaced by human review if needed. We use Claude Code as an interactive interface for our team. It queries our central intelligence via GraphQL. The intelligence layer does not depend on Claude Code. ## Why not fully sovereign? We could run the entire pipeline on open-weight models. Replace Haiku triage with a fourth Scaleway model. Replace Sonnet arbitration with human review only. We chose not to — for now. The Anthropic models genuinely improve quality at the triage and arbitration stages. They are better at nuanced evaluation than any open-weight model currently available to us. Removing them would mean more false positives reaching our training system, or more human review time. The key is that this is a choice, not a dependency. The architecture supports full sovereignty. We opt into Anthropic where the quality justifies it, and we can opt out at any time without redesigning the system. This is the same principle we advocate to our clients: freedom to operate means your systems work for you, you can change providers when you want, and you stay because you choose to — not because you are locked in. ## The data store PostgreSQL with pgvector on Kubernetes, managed by CloudNativePG. Hasura generates a typed GraphQL API from the database schema. All on Scaleway, all European infrastructure. Backups to Scaleway Object Storage. The database holds the structured knowledge base: insights with confidence scores, buyer personas, case studies, competitive positioning. The training system queries this database to generate exercises for our sales team. ## What we learned **Use a different model to judge than to create.** Self-evaluation is unreliable. Three independent judges catch what self-scoring misses. **Field validation is not optional.** LLMs generate creative field names. One run produced "refidence" instead of "evidence." Strict schema validation before database insertion is essential. **Replace, do not accumulate.** Running the same extraction twice produces slightly different results. Each ingestion run replaces all records from that source. The database always reflects the latest extraction, not an accumulation of probabilistic runs. **Design for graceful degradation.** Every external dependency should have a fallback. Our fallback for Anthropic is human review. Our fallback for Scaleway is self-hosting the same open-weight models. Our fallback for any single scoring model is the other two. **Sovereignty is a spectrum, not a binary.** Fully sovereign is possible but has a quality cost. Strategically using non-European services in non-critical roles — with clear fallback paths — is a pragmatic approach that most organizations can adopt today. --- ## Retention Theatre or how to Engage your Developer Community URL: https://clouds-of-europe.eu/content/strategy-transition/leadership-guides/retention-theatre-or-how-to-engage-your-developer-community Author: Jurg van Vliet Published: 2026-03-12 Category: Strategy & Transition Type: Leadership Guides The honest answer is that perks like catered meals, gym memberships, and office goodies are **retention theater**. They're visible, they're photographable for the careers page, and they feel generous. But they operate at the wrong layer of the retention problem. They address comfort, not meaning. An engineer who is bored, blocked, or underutilized doesn't become engaged because lunch is free. They just eat better while quietly job hunting. **The research on this is fairly consistent:** extrinsic perks have a short hedonic half-life. They become expected within weeks, stop registering as positive, and only become salient again if removed — at which point they're a reason to leave. You're essentially paying a maintenance cost on something that was never driving retention in the first place. **A good developer experience compounds differently.** Every hour of friction removed is an hour of engagement added, permanently. Every abstraction that lets an engineer ship faster raises their baseline satisfaction with the work itself. It improves the thing they actually care about — their craft. The spending comparison is stark when you frame it this way. A catered lunch program for 15 people might run €50-80k per year. That same budget, invested in platform engineering or tooling, produces something that improves daily experience indefinitely and — in your case specifically — becomes a client-facing asset. **The one thing perks do well** is signal culture during recruiting. A nice office and good food say "we take care of people here." But that signal needs to be backed by the actual experience of working there, or it becomes a source of cynicism very quickly. The companies that retain developers well tend to spend less on visible perks and more on invisible ones — fast CI, good hardware, clear processes, time to do things properly. Engineers notice the invisible ones more, because they feel them every day. --- ## The Shopping List for Europe's Core Cloud Providers URL: https://clouds-of-europe.eu/content/policy-sovereignty/provider-ecosystem/the-shopping-list-for-europe-s-core-cloud-providers Author: Jurg van Vliet Published: 2026-03-10 Category: Policy & Sovereignty Type: Provider Ecosystem Tags: cloudprovider, europeancloud, iaas, infrastructure, kubernetes, sovereignty Europe will have three to five general-purpose cloud providers with pan-European presence and true scale. Not twenty. Not one. A small number of providers — rooted in France, Germany, the Nordics, Eastern Europe, and perhaps a fifth entrant — each operating ten or more regions worldwide. They will be distinctly European in ownership, governance, and values. Some will become global powerhouses. An organisation building its cloud backbone will pick one or two of them. Behind this handful of core providers, dozens of niche and specialty clouds will thrive — Kubernetes-first platforms, green computing, sovereign hosting for regulated industries, GPU farms for AI workloads. That ecosystem matters. But the backbone needs breadth. ## Why Kubernetes-Native Changes the Game Most cloud comparisons dwell on managed databases, AI services, and serverless functions. That misses the point. When you run Kubernetes-native, you bring your own data layer, your own observability, your own service mesh. The provider supplies infrastructure primitives — compute, network, storage — and stays out of the way. This matters because it reframes the competitive landscape — and the economics. Hyperscalers charge a premium on commodities — compute, storage, network — because those margins subsidise their managed services. When you replace RDS with CloudNativePG, you stop paying for the service but you still pay the inflated commodity price. A provider that does not offer managed services does not need to recover that cost. The numbers bear this out: a comparable Kubernetes setup on Scaleway costs 65% less than AWS, and on STACKIT 57% less. That gap is not efficiency — it is business model. And it is very hard for a hyperscaler to change. Repricing commodities would cannibalise the revenue that funds their service catalogue. European providers, unburdened by that catalogue, can price infrastructure for what it costs. The gap between a managed service and a Kubernetes operator is trust. Once organisations trust the operator, the hyperscaler's service catalogue stops being a moat — and its pricing becomes a tax. Yes, this filters out enterprises still running VMware and SAP on bare metal. Intentionally. This article is for organisations that have made the Kubernetes-native choice — or are about to. ## Europe Can Lead, Not Just Catch Up The temptation is to frame European cloud as a follower trying to replicate what AWS built. That is the wrong frame. Europe can define the next generation of cloud infrastructure. **Regulation as product.** GDPR, NIS2, DORA, the EU AI Act — European providers live inside this regulatory reality. Instead of treating compliance as overhead, they can embed it as a platform feature: data residency as a primitive, audit trails as infrastructure, sovereignty as a toggle. No hyperscaler can do this natively because their architecture spans jurisdictions by design. **Open source as moat.** CloudNativePG, Cilium, Flux — the critical Kubernetes infrastructure stack is European-led. Investment in open source creates ecosystem lock-in without vendor lock-in. That is the opposite of the hyperscaler model, and it is Europe's single greatest strategic advantage in cloud. **Sustainability as standard.** European energy regulations and carbon pricing give EU providers a structural incentive to build efficient infrastructure. Waste heat reuse, renewable-powered data centres, and energy-proportional computing are niche curiosities today. They will be competitive requirements globally. **Federation over centralisation.** What if European providers did not each need to build the full stack alone? Exoscale's eight zones across six countries paired with Scaleway's three-AZ regions through a standard interconnect layer would be more powerful than either alone. This is what Gaia-X promised but failed to deliver technically. The opportunity remains. ## The Hard Floor: Regions and Zones At least three EU regions. At least three availability zones per region. Independent power, cooling, and networking in each zone. This is the minimum topology for a resilient, sovereign deployment. Every successful public cloud — AWS, Azure, GCP — converged on this architecture, and it is default Kubernetes. Without it, you cannot place workloads close to users across the continent. You cannot survive a zone failure without downtime. You cannot run synchronous database replication across zones — which needs sub-2ms latency, only achievable within a region. Every other requirement builds on this foundation. Between regions, expect 5-10ms latency for close neighbours (Frankfurt to Amsterdam around 8ms) and 10-25ms for more distant pairs (Paris to Stockholm, Frankfurt to Warsaw). Synchronous replication across regions is impractical; asynchronous replication with well-designed failover is the pattern. Today, only one European provider meets this bar fully: Scaleway, with three regions (Paris, Amsterdam, Warsaw), each offering three availability zones. OVHcloud has the geographic spread but only Paris has three zones. STACKIT has three zones in Germany but only two regions, focused on DACH. The gap is real. ## P0 — Non-Negotiable These requirements define whether a provider qualifies at all. **Managed Kubernetes.** Upstream CNCF-conformant. Control plane spans availability zones with HA guarantees. No restrictions on Custom Resource Definitions — operators like CloudNativePG, cert-manager, and Flux must install and run without interference. The CNI must support Kubernetes NetworkPolicy (Cilium preferred). No blocking or throttling of admission webhooks. Sufficient etcd capacity for GitOps-scale CRD counts. Latest stable Kubernetes available within 30 days of upstream release, with n-2 version support and managed rolling upgrades with rollback. **Compute.** Diverse instance types: general purpose, compute-optimised, memory-optimised, and GPU. ARM nodes for cost and energy efficiency on stateless workloads. **Networking.** VPC with private subnets for full tenant isolation. Cross-region VPC peering so clusters talk privately without touching the public internet. Cloud-native load balancers (L4 and L7) integrated with Kubernetes Service and Ingress. Managed authoritative DNS with health checks and failover. DDoS protection at the network level, included — not an upsell. Private interconnect for dedicated links to on-premise or other providers. **Storage.** Block storage with a CSI driver — SSD-backed, with volume snapshots and online expansion. Topology-aware storage classes so pods and their volumes land in the same zone. S3-compatible object storage for backups, artifacts, and WAL archiving. Encryption at rest with customer-managed keys. **Identity.** OIDC integration for kubectl, CI/CD pipelines, and service accounts. Scoped IAM per cloud resource — separate keys for storage, registry, and DNS, not a single god credential. **Operations.** API-first: a complete Terraform or OpenTofu provider, plus REST APIs for every resource. If you cannot codify it, you cannot operate it at scale. Exportable audit logging for API calls and authentication events. No egress fees between zones, and low inter-region rates. Transparent control plane pricing — free or predictable, not $73 per month per cluster. ## P1 — Expected at Scale These matter as you grow from one cluster to ten, and from one team to many. - Private OCI-compliant container registry, per-project, with pull credentials integrated into the cluster - IPv4 and IPv6 dual-stack — native, not bolted on - Pod Security Standards with enforce and audit modes - RBAC mapped to an external identity provider - Automated service account key rotation - Programmatic billing API per project, cluster, and namespace ## P2 — Differentiators Nice to have. Often solved at the platform layer rather than by the provider. - Node autoscaling with scale-to-zero and sub-two-minute node readiness - Spot and preemptible instances for batch and CI workloads - Bare metal nodes for latency-sensitive or hardware-isolated workloads - Bring-your-own-IP for BGP anycast or IP reputation continuity - Multi-attach (RWX) volumes for shared storage - Geo-replicated container registry - Built-in vulnerability scanning and image retention policies ## The Scorecard We assessed all European providers we could find that offer managed Kubernetes and general-purpose IaaS. Here is where they stand against the hard requirements. ### The Contenders | Provider | HQ | Regions | 3-AZ Regions | Managed K8s | Cross-Region VPC | IaC Coverage | |----------|:--:|:-------:|:------------:|:-----------:|:----------------:|:------------:| | **Scaleway** | FR | 3 (PAR, AMS, WAW) | 3 of 3 | CNCF | No | Good | | **OVHcloud** | FR | 15+ | 1 (Paris) | CNCF | vRack (L2) | Partial | | **STACKIT** | DE | 2 (DE, AT) | 1 (DE) | CNCF | Unverified | Good | | **Open Telekom Cloud** | DE | 2 (DE, NL) | 1 (DE) | Yes | Yes | Good | | **Ionos** | DE | 4 (DE, ES, UK) | Partial | Yes | Limited | Partial | | **Exoscale** | CH | 8 zones, 6 countries | 0 | CNCF | No | Good | | **UpCloud** | FI | 10 EU locations | Unverified | Yes | Unverified | Partial | **Scaleway** leads on the hard floor: three regions, three zones each, CNCF-certified Kubernetes, strong Terraform coverage, free mutualized control plane. The gap: no cross-region VPC peering. Germany expansion is announced. **OVHcloud** has the broadest geographic reach and anti-DDoS included by default. Cross-region networking works through vRack (L2 extension). But only Paris qualifies as a true 3-AZ region, and operational maturity on Kubernetes lags behind the marketing. The Strasbourg fire of 2021 is a reminder that geographic spread means nothing without proper zone isolation. **STACKIT** is backed by Schwarz Group (Lidl/Kaufland), giving it financial depth. Three zones in Germany, but the Austria region is still building out. DACH-focused — no pan-European ambition yet. **Exoscale** takes an interesting approach: eight zones across six countries (Germany, Austria, Switzerland, Bulgaria, Croatia), each a standalone location. Great geographic diversity, but no multi-AZ regions means no zone-level HA within a single cluster. Swiss jurisdiction is a sovereignty advantage. Notably, Exoscale appears to be among the first European providers — possibly the first — to integrate Karpenter as a native add-on, with scale-to-zero and sub-minute node provisioning included in its managed Kubernetes offering. For workloads that do not need multi-AZ resilience (batch, CI/CD, GPU inference), this is smart positioning: resource-conscious infrastructure for a resource-conscious continent. In a federated model, Exoscale's footprint becomes very attractive. **Open Telekom Cloud** has strong enterprise credentials via Deutsche Telekom, but is built on Huawei technology — a sovereignty concern that cuts both ways. Chinese cloud providers entering the European market with local subsidiaries and data residency is a scenario that merits attention, not dismissal. ### The Niche Players These providers serve specific markets well but lack the breadth or geographic reach for a pan-European backbone: | Provider | HQ | Focus | |----------|:--:|-------| | **Hetzner** | DE | Price leader, no native managed K8s | | **Civo** | UK | Kubernetes-first, 90-second clusters | | **Infomaniak** | CH | Swiss sovereignty, K8s launched Jan 2026 | | **plusserver** | DE | Gaia-X / Sovereign Cloud Stack | | **Cleura** | SE | Regulated industries, OpenStack-based | | **Elastx** | SE | Stockholm 3-AZ, OpenStack-based | | **SysEleven** | DE | MetaKube, German market | | **gridscale** | DE | DE/NL/AT/CH, mid-market | | **Leafcloud** | NL | Green computing, waste heat reuse | ## The Coming of Age Moment There is a black swan that would accelerate everything: a US executive order or trade action restricting data flows to European cloud services. The Privacy Shield was already invalidated. The Data Privacy Framework is politically fragile. If it falls, European organisations will need sovereign alternatives overnight — and the providers who meet the requirements in this article will see demand spike faster than they can build capacity. This is not a threat. It is the defining moment for the European cloud ecosystem. The question is not whether it will happen, but whether European providers will be ready when it does. Some will argue that a hyperscaler could simply create a legally separate European entity and achieve sovereignty certification. Europeans have heard this before. They are tired of subsidising structures designed to circumvent the intent of their own regulations. The market is moving past "sovereign enough for the auditor" toward genuine European ownership and control. ## What This Means An organisation does not need a hyperscaler's service catalogue. It needs three regions, three zones, solid networking, and an API for everything. The race is on. Scaleway is ahead on architecture. OVHcloud has the data centres but needs to build out zones. STACKIT needs to look beyond DACH. The providers that close these gaps first will anchor Europe's cloud backbone. The ones that do not will remain — valuable, but niche. The provider that delivers this — sovereign, performant, and boring — wins. #kubernetes #iaas #sovereignty #cloudprovider #infrastructure #europeancloud --- ## Software That (Almost) Runs Itself URL: https://clouds-of-europe.eu/content/policy-sovereignty/open-source-standards/software-that-almost-runs-itself Author: Jurg van Vliet Published: 2026-03-09 Category: Policy & Sovereignty Type: Open Source & Standards ## Source code and tribal knowledge In the beginning, software shipped as source code. You downloaded a tarball, ran `./configure && make && make install`, and hoped your system had the right libraries. The Makefile encoded how to *build* the software. Everything after that — where to put the config file, how to start the daemon on boot, what to do at 3 AM when the process died — lived in a single place: the sysadmin's head. The sysadmin was the expert. She tuned kernel parameters, wrote init scripts, memorized the order in which services had to start. Her knowledge was deep, specific, and almost entirely undocumented. When she left, the knowledge left with her. Package managers — RPM in 1997, dpkg before that — improved distribution by solving dependency resolution and file placement. But they didn't touch the harder problem: *how do I run this thing reliably?* That stayed tribal. Post-mortems were written, wikis were updated, and nobody read either of them twice. ## Portable but not operational Then, in August 2006, Amazon launched EC2. The unit of deployment shifted from "package on a server" to "entire machine image." AMIs captured not just the application but its runtime environment. When Docker followed in March 2013, it sharpened the idea further: a container image was a self-contained, portable artefact. No more "works on my machine." The application carried its dependencies with it. This was a genuine revolution in *distribution*. But portable is not the same as operational. A container image tells you nothing about how many replicas to run, how to route traffic between them, or what to do when one crashes. The complexity didn't disappear. It migrated — from "how do I install this" to "how do I orchestrate this at scale." The human role shifted with it. In October 2009, Patrick Debois organized the first DevOpsDays conference in Ghent, giving a name to what many teams were already feeling: developers couldn't just throw code over the wall to operations anymore. The cloud had blurred the boundary. SaaS products solved the problem entirely for end users — someone else runs the infrastructure. But for the teams *building* those platforms, operational burden compounded with every microservice added to the architecture. The sysadmin became the DevOps engineer: someone who understood CI/CD pipelines, container orchestration, and monitoring, not just the OS. ## Infrastructure as code, application as mystery The next leap came in February 2011, when AWS released CloudFormation. For the first time, you could describe infrastructure — servers, networks, load balancers — as a declarative template. Terraform, released by HashiCorp in July 2014, generalized the idea across cloud providers and quickly became the industry standard. (Its open-source fork OpenTofu now continues that work.) Infrastructure became code: reviewable, versioned, reproducible. The DevOps engineer became the cloud engineer — someone who thought in resource graphs and state files rather than SSH sessions and shell scripts. But Infrastructure as Code has a blind spot, and it's a fundamental one. It describes what infrastructure should *exist*: this many instances, this network topology, this database cluster. It does not describe how the applications running on that infrastructure should *behave*. Terraform can provision a PostgreSQL cluster. It cannot promote a replica when the primary fails. It can create Kubernetes nodes. It cannot safely restart a stateful service that holds time-series data in memory. Infrastructure as Code answers "what should exist." It is silent on "what should happen next." --- ## The complexity cliff That silence becomes deafening as software grows more complex. Consider a production deployment of Grafana Mimir, a horizontally scalable metrics backend. It runs as eight distinct components: distributors, ingesters, queriers, query-frontends, query-schedulers, compactors, store-gateways, and rulers. Each component scales independently. Each has different failure modes. The ingester holds recent metric data in memory, organized by tenant. Roll it carelessly during an upgrade — kill the pod before it flushes to long-term storage — and those samples vanish. The store-gateway needs time to load blocks from object storage before it can serve queries. Restart it too early and dashboards show gaps. No Makefile builds this operational understanding. No container image encodes it. No Terraform module captures the safe order in which to restart eight interdependent components across three availability zones. The knowledge lives, as it always has, in human heads — and in runbooks that grow stale faster than anyone can maintain them. This is the gap the Kubernetes operator pattern fills. --- ## Ship the expertise with the software In November 2016, engineers at CoreOS published a blog post titled "Introducing Operators: Putting Operational Knowledge into Software." The insight was direct: the best people to automate a database are the people who built the database. They know the failure modes. They know the recovery sequences. They know which upgrade paths are safe and which corrupt data. Instead of writing all that down and hoping someone follows the instructions, encode it in a program that runs alongside the application. Ship the expertise with the software. An operator is a controller — a long-running process in the Kubernetes cluster that watches the desired state of a resource, compares it to reality, and reconciles the difference. This sounds like Terraform, but the distinction matters. Terraform runs when you invoke it, then stops. An operator runs *continuously*. It doesn't just create infrastructure; it reacts to drift, failure, and change in real time. Grafana Labs ships a concrete example of this idea with their Mimir Helm chart: the rollout-operator. Its scope is narrow — it manages the safe rollout of stateful components during upgrades in zone-aware deployments — but within that scope, it replaces an entire runbook. When Mimir runs with zone-aware replication across three availability zones, an upgrade must proceed one zone at a time. Update zone A, wait for those ingesters to become ready and catch up on replication, then move to zone B, then zone C. At no point should two zones degrade simultaneously. The rollout-operator coordinates this sequence automatically. It defines two custom resource types — `ZoneAwarePodDisruptionBudgets` and `ReplicaTemplates` — that extend the Kubernetes API with Mimir-specific operational concepts. The budget ensures zone-safe disruption constraints. The template maintains consistent replica counts per zone during the transition. A human operator could do all of this by hand. Read the docs. Run kubectl commands in the right order. Watch logs. Wait. Proceed. But the human does it once per upgrade, under pressure, possibly at 2 AM, and gets it wrong one time in twenty. The rollout-operator does it every time, identically, because it *is* the procedure. --- ## Knowledge that compounds The pattern has spread far beyond monitoring backends. CloudNativePG encodes PostgreSQL DBA knowledge: when a streaming replica falls behind, it promotes a new one, reconfigures replication, and updates service endpoints — the same decision tree a senior DBA would follow, executed in seconds rather than the minutes it takes a human to context-switch and assess the situation. The Istio operator manages service mesh configuration. Strimzi manages Kafka clusters. The CNCF Operator Framework and Kubebuilder have made it straightforward to build new ones, and Helm charts increasingly bundle operators as dependencies. This changes what gets distributed. Software is no longer just an image and documentation. It's an image, an operator, and a set of custom resource definitions that extend Kubernetes with domain-specific vocabulary. When you install Mimir, you don't just get containers. You get new API objects that represent Mimir's operational model. The platform becomes aware of concepts — zones, replicas, rollout safety — that it previously had no language for. The compounding effect is the most significant consequence. When the CloudNativePG team encounters a new failure mode in production, the fix becomes a patch to the operator's logic. It ships to every deployment that upgrades, not just the cluster that burned. Operational knowledge accumulates in code, not in runbooks that age into fiction. --- ## The platform engineer The human role has shifted once more. The **platform engineer** — the title gaining traction now — doesn't manually configure Mimir or write custom rolling-update scripts. She selects operators, configures their custom resources, and trusts them to encode the expertise she once carried herself. Her job is to compose the right operators, define the right constraints, and intervene when automation reaches its limits. Each era eliminated a category of manual work and demanded a different kind of expertise. The sysadmin knew machines. The DevOps engineer knew pipelines. The cloud engineer knew resource graphs. The platform engineer knows operators. The expertise didn't diminish — it shifted to a higher level of abstraction. --- ## The cost of delegation Operators are not free, and it would be dishonest to pretend otherwise. They add moving parts. An operator bug can cause more damage than a manual mistake, precisely because it acts faster and at broader scale. An operator that loses its watch connection to the Kubernetes API may misread the current state and make destructive decisions with full confidence. The model concentrates trust. When the rollout-operator decides the order in which to restart your ingesters, you trust that Grafana Labs tested that sequence against your topology. They usually did. But sometimes your deployment is the edge case nobody anticipated. And the complexity doesn't vanish — it transforms. You no longer need to know how to manually coordinate a zone-by-zone rollout. But you need to understand custom resource definitions, admission webhooks, RBAC permissions, and the failure modes of the operator itself. The old complexity was chaotic and human-dependent. The new complexity is structured and machine-dependent. Both can fail you. --- ## Where this goes Still, the trajectory is clear. Software distribution converges on a model where the artefact and its operational knowledge ship as one unit. The DBA's expertise lives in the PostgreSQL operator. The SRE's upgrade checklist lives in the rollout-operator. The network engineer's traffic-shifting logic lives in the service mesh controller. This doesn't eliminate human judgment. Someone must still decide *what* to run, *why* to run it, and *when* the automated decisions are wrong. But the gap between "I want to run this" and "this runs correctly in production" narrows with every operator that ships. The deepest shift may be cultural. When software knows how to run itself, the vendor's responsibility extends beyond the artefact. Shipping a container image without an operator starts to feel like shipping a car without a transmission — technically complete, practically useless for most drivers. The best infrastructure teams already think this way. The rest will follow, because the alternative — shipping ever-growing complexity without encoding the expertise to manage it — stopped scaling years ago. --- ## Qualtrics' Marketing Plays You for a Fool URL: https://clouds-of-europe.eu/content/policy-sovereignty/policy-regulation/qualtrics-marketing-plays-you-for-a-fool Author: Jurg van Vliet Published: 2026-03-04 Category: Policy & Sovereignty Type: Policy & Regulation It's already happening in the Netherlands. Dutch universities are cutting Qualtrics access or dropping it entirely. The reason is straightforward: Qualtrics switched from per-institution licensing to per-user and per-response pricing models. For some universities, costs doubled or tripled overnight. In a sector facing millions in budget cuts, a survey tool that suddenly costs tens of thousands more per year is an easy target. The University of Twente switched to Crowdtech Survey in 2025. Other institutions are restricting Qualtrics to specific faculty research — students who previously had unlimited access are locked out. The tool that was supposed to be campus-wide infrastructure is becoming a luxury line item. This isn't a Dutch problem. It's global. Duke University in the US renewed their Qualtrics contract at 2.5 times their 2023 price — and that was the *negotiated* outcome. The initial proposal from Qualtrics? A 6x increase over the contract term. The University of Georgia dropped Qualtrics for QuestionPro because they could no longer justify it as an enterprise-wide tool. Qualtrics applies a minimum 5% annual uplift on all renewals and will not negotiate it away, regardless of how much your usage grows. This is what happens when a vendor pivots from academia to "Experience Management" for the commercial sector. You're no longer the customer. You're the customer they're extracting value from on the way out. And while your procurement team scrambles to renegotiate, Qualtrics' marketing team is in Seattle this month — X4 2026 — announcing AI features you didn't ask for and can't afford. ## The GDPR Distraction Qualtrics markets GDPR compliance hard. Their website has a dedicated GDPR page. They offer Data Processing Agreements, EU Standard Contractual Clauses, data subject rights tooling, and pseudonymisation features. All of which is theatre. Not because Qualtrics doesn't implement these measures — they do. But because GDPR compliance does not solve the actual legal problem European universities face. The problem is the CLOUD Act. Qualtrics is a US company. Under the US CLOUD Act (2018), US authorities can compel any US-headquartered company to hand over data stored anywhere in the world. It doesn't matter if the servers are in Frankfurt. It doesn't matter if the DPA says "EU-only." When a US court issues a CLOUD Act order, Qualtrics must comply. And they cannot tell you about it. European data protection authorities have been clear: the CLOUD Act is not a sufficient legal basis under GDPR for transmitting personal data to US authorities. An organisation processing data through a US platform subject to CLOUD Act jurisdiction may be in ongoing structural breach — not because of a specific incident, but because of the architecture itself. Your student survey data. Your employee satisfaction research. Your longitudinal studies on mental health. All of it — structurally exposed to US government access, with no legal mechanism to prevent it. Qualtrics knows this. That's why they talk about GDPR so loudly. It's misdirection. They want you debating Data Processing Agreements while ignoring the jurisdiction of the company holding your data. ## Functionality Is Not the Moat They Think It Is Qualtrics built its market position when open source survey tools were primitive. That was a decade ago. The landscape has changed. Formbricks — a German open source survey platform licensed under AGPLv3 — now provides enterprise-grade survey capabilities: link surveys, in-app surveys, event-triggered surveys, NPS, CSAT, full API access, and native integrations with Slack, Google Sheets, Airtable, and automation platforms like n8n and Make.com. But here's what matters for research institutions: Formbricks is one component of a complete research platform. Pair it with PostgreSQL for central data storage, JupyterHub for R and Python analysis, and Apache Superset for dashboards and visualisation — and you have a stack that doesn't just match Qualtrics. It exceeds it. Why? Because researchers don't just need surveys. They need analysis. Qualtrics forces you into their analysis tools, their export formats, their workflow. An open source stack gives researchers JupyterHub notebooks — the same environment they already use for statistical analysis — connected directly to their survey data. No export. No format conversion. No waiting for Qualtrics to add the R package you need. Every component is open source. Every component runs on Kubernetes. Every component deploys with Helm charts. You can audit the code, extend the platform, and leave without losing your data. Try that with Qualtrics. ## The Cost Comparison Is Embarrassing Qualtrics doesn't publish pricing. That's intentional. Opacity is a pricing strategy. When you can't compare, you can't negotiate. So let's compare anyway. A university-scale Qualtrics deployment runs €80,000–€150,000 per year. Over five years, that's €400,000–€750,000. And remember: Duke's renewal went up 2.5x. Your number is going in one direction. An open source research platform at university scale — managed cloud, unlimited researchers, multi-faculty isolation, 99.9% SLA — runs €60,000–€100,000 per year. On-premise with annual maintenance: €36,000–€50,000 per year after initial setup. Five-year total: €300,000–€500,000. But the real comparison is worse for Qualtrics, because the open source platform includes capabilities Qualtrics charges extra for or doesn't offer at all: - **JupyterHub** for multi-user R/Python analysis — included, not a separate license - **Apache Superset** for BI dashboards — included, not a third-party add-on - **Full data export** in open formats (CSV, JSON, SPSS-compatible) — always, not as a premium feature - **Source code access** — audit, extend, or walk away. No lock-in - **No per-seat licensing** — pricing by organisational scope, not headcount The pricing model itself is the tell. Qualtrics prices by user, by response, by feature tier — designed to extract maximum revenue as usage grows. An open source alternative prices by organisational scope with transparent consumption rates. You know what you'll pay. You can plan for it. ## What This Is Really About This isn't about Qualtrics being a bad product. It works. Researchers know how to use it. The surveys look professional. This is about a European institution sending student data, employee feedback, and research results to a US company that is legally compelled to hand that data to US authorities on request — while paying €100,000+ per year for the privilege — while an open source alternative built in Germany, hosted in Europe, running on your infrastructure, exists at lower cost with comparable functionality. The University of Twente figured this out. They moved to an alternative. Other Dutch universities are restricting access, quietly acknowledging that the value equation no longer holds. Students who built their thesis research around Qualtrics now find themselves locked out because the institution can't justify the per-user cost. Every euro spent on Qualtrics is a euro that doesn't go to a researcher. Every survey response stored in Qualtrics is a data point under US jurisdiction. Every year the contract renews is a year closer to the next price increase you can't negotiate. Qualtrics' marketing team is paid to make you forget all of this. They'll talk about AI. They'll talk about experience management. They'll talk about GDPR compliance. They will not talk about the CLOUD Act. They will not publish their pricing. They will not give you the source code. They will not let you run it on your own infrastructure. That tells you everything you need to know. ## What You Can Do If you're at a European university evaluating Qualtrics renewal: 1. **Ask Qualtrics directly:** "Can you guarantee that our data will never be subject to a CLOUD Act order?" They can't. Watch how they answer. 2. **Request their pricing in writing** before negotiation. Compare it to published, transparent alternatives. Notice the difference. 3. **Evaluate Formbricks + JupyterHub + Superset** on a single department. Run it for one semester. Let your researchers compare. The results speak for themselves. 4. **Talk to your DPO.** Ask them about structural CLOUD Act exposure. Not the GDPR theatre — the actual jurisdictional risk of using a US platform for European research data. The tools exist. The cost is lower. The data stays in Europe. The source code is open. The only thing Qualtrics has left is your switching cost — and their marketing budget. Don't let either play you for a fool. --- *This article is published on [Clouds of Europe](https://clouds-of-europe.eu), a practitioner community building European cloud independence. We're not selling an alternative — we're documenting that one exists.* --- **Sources:** - [Duke University Qualtrics contract renewal — 2.5x price increase](https://oit.duke.edu/news/duke-renews-contract-qualtrics/) - [Qualtrics 5% minimum renewal uplift — Vendr negotiation data](https://www.vendr.com/marketplace/qualtrics) - [CLOUD Act vs. GDPR — structural conflict analysis](https://www.exoscale.com/blog/cloudact-vs-gdpr/) - [CLOUD Act risk for European data](https://xpert.digital/en/us-cloud-act/) - [Formbricks — open source survey platform (Germany)](https://formbricks.com/) - [Qualtrics GDPR compliance page](https://www.qualtrics.com/gdpr/) - [University of Twente switches to Crowdtech Survey (2025)](https://www.utwente.nl) - [University of Georgia drops Qualtrics for QuestionPro](https://blog.uxtweak.com/qualtrics-pricing/) - [Negative 2026 outlook for higher education sector](https://www.quantumrun.com/consulting/higher-education-costs/) --- ## Your Go Builds Phone Home to Google. Every Single Time. URL: https://clouds-of-europe.eu/content/policy-sovereignty/open-source-standards/your-go-builds-phone-home-to-google-every-single-time Author: Jurg van Vliet Published: 2026-03-03 Category: Policy & Sovereignty Type: Open Source & Standards Every `go build` sends your organisation's build metadata to Google. Module names, versions, timestamps, your CI runner's IP address. Google retains those IPs for [30 days](https://proxy.golang.org/privacy). A mid-size engineering team runs a hundred Go builds a day — each fetching dozens of modules. That is thousands of requests, every working day, carrying your technology choices to a US corporation subject to FISA Section 702. ## What the metadata reveals A Go build fetches dozens of modules. Each fetch is an HTTP request to `proxy.golang.org` carrying the module path, the requested version, and the source IP. Aggregate that over weeks and the picture sharpens: - **Technology stack.** Module paths reveal your dependencies: which web framework, which database driver, which cloud SDK. A competitor or adversary learns what you build on. - **Release cadence.** Version bumps in `go.sum` trigger fresh fetches. The pattern of fetches over time reveals how often you ship. - **Team structure.** Different IP ranges fetching different module sets map to different teams and projects. - **Vendor relationships.** Fetching `cloud.google.com/go`, `github.com/aws/aws-sdk-go`, or `github.com/Azure/azure-sdk-for-go` signals which cloud providers you use and when you migrate. None of this is classified. All of it is commercially sensitive. And all of it lands on infrastructure controlled by a company that the Court of Justice of the European Union [ruled](https://curia.europa.eu/juris/liste.jsf?num=C-311/18) cannot guarantee adequate protection under EU law. ## The legal problem nobody budgets for Schrems II (2020) invalidated Privacy Shield because US surveillance law — specifically FISA Section 702 — gives US authorities access to data held by US companies, without meaningful judicial oversight for non-US persons. The fundamental incompatibility between FISA 702 and GDPR Articles 44-49 has not been resolved. The EU-US Data Privacy Framework (2023) faces the [same structural challenge](https://noyb.eu/en/european-commission-gives-eu-us-data-transfers-third-round-cjeu). For most SaaS tools, companies manage this with DPAs and Standard Contractual Clauses. For a build chain, nobody even thinks about it. There is no DPA with Google for proxy.golang.org. There is no contractual basis for the transfer. Your CI pipeline sends metadata to a US company hundreds of times a day and the legal basis is "we didn't notice." A CTO at a regulated European company — financial services, healthcare, public sector — has a freedom-to-operate problem hiding in plain sight. Not because proxy.golang.org is malicious, but because the data transfer has no legal footing and the exposure is continuous. ## The fix is one environment variable ``` export GOPROXY=https://goproxy.eu,https://proxy.golang.org,direct ``` [goproxy.eu](https://goproxy.eu) is a free, community-operated Go module proxy running on [Scaleway](https://www.scaleway.com/) in Paris, Amsterdam, and Warsaw. It caches modules regionally. Once a module is cached, subsequent fetches from that region never reach Google. What changes: - **Popular modules** (the vast majority of fetches) serve from European cache. No request to Google. - **Cold modules** fall through to proxy.golang.org on first fetch, then cache regionally for everyone. - **No IP logging.** Client IPs are stripped from access logs at the Envoy Gateway level and regex-redacted from all log pipelines before storage. They never reach our observability stack. - **EU jurisdiction only.** Scaleway is a French company (Iliad Group). No CLOUD Act exposure, no FISA 702. The fallback chain means you lose nothing. If goproxy.eu is down, your toolchain falls through to Google's proxy, then to direct VCS access. Builds never break. ## For CI/CD at scale A typical `go build` fetches 20-80 modules. A hundred builds a day — normal for a team of 15-20 Go developers with CI on every push — means 2,000-8,000 module requests per working day. After a week of cache warming, 90%+ serve from European cache. The remaining cold fetches — new dependencies, version bumps — still fall through to Google, but the metadata exposure drops from "everything, every time" to "occasional new modules." GitLab CI: ```yaml variables: GOPROXY: "https://goproxy.eu,https://proxy.golang.org,direct" ``` GitHub Actions: ```yaml env: GOPROXY: "https://goproxy.eu,https://proxy.golang.org,direct" ``` Dockerfile: ```dockerfile ENV GOPROXY=https://goproxy.eu,https://proxy.golang.org,direct ``` No registration. No API keys. No authentication. One variable, all pipelines. ## If this isn't enough, run it yourself goproxy.eu strips IPs and runs on European infrastructure, but it is still someone else's server. If your threat model requires full control, the entire stack is open source: OpenTofu modules, Flux GitOps manifests, Envoy Gateway configuration, the Alloy log pipeline with IP redaction, alert rules, dashboards. Same code that runs in production. [Source repository](https://gitlab.aknostic.com/aknostic/goproxy) | [Architecture](https://gitlab.aknostic.com/aknostic/goproxy/-/blob/main/docs/architecture.md) The second article in this series covers the full engineering breakdown: architecture, cost per region, operational trade-offs, and what we'd do differently. --- *goproxy.eu is a [Clouds of Europe](https://clouds-of-europe.eu) community project.* --- ## European Go Proxy for EUR 93 per Month URL: https://clouds-of-europe.eu/content/practice/implementation-patterns/european-go-proxy-for-eur-93-per-month Author: Jurg van Vliet Published: 2026-03-03 Category: Practice Type: Implementation Patterns [goproxy.eu](https://goproxy.eu) is a multi-region Go module proxy serving European developers from Paris, Amsterdam, and Warsaw. Total infrastructure cost: EUR 93 per month. No proprietary dependencies. No vendor lock-in. Every line of configuration is in [a public Git repository](https://gitlab.aknostic.com/aknostic/goproxy). This article is the full engineering teardown. Architecture decisions, cost breakdown, what we tried and rejected, and operational numbers after running in production. ## Architecture in 30 seconds Three Kubernetes clusters on [Scaleway](https://www.scaleway.com/) Kapsule (Paris, Amsterdam, Warsaw). Each runs [Athens](https://github.com/gomods/athens) behind [Envoy Gateway](https://gateway.envoyproxy.io/) with TLS from Let's Encrypt. Module cache in regional S3-compatible Object Storage. Redis for SingleFlight deduplication. [Scaleway GeoDNS](https://www.scaleway.com/en/docs/domains-and-dns/how-to/manage-dns-records/) routes requests to the nearest region. [Flux CD](https://fluxcd.io/) reconciles everything from Git. ``` GOPROXY=https://goproxy.eu,https://proxy.golang.org,direct goproxy.eu (GeoDNS ALIAS) │ ┌──────┼──────┐ │ │ │ fr-par nl-ams pl-waw ← geo IP routes to nearest │ │ │ Envoy Envoy Envoy ← TLS, rate limiting (1000 req/min) │ │ │ Athens Athens Athens ← Go module proxy │ │ │ Redis Redis Redis ← SingleFlight dedup │ │ │ S3 S3 S3 ← regional cache │ │ │ └──────┼──────┘ │ proxy.golang.org ← upstream on cache miss ``` ## Cost breakdown | Component | Per region | 3 regions | |-----------|-----------|-----------| | Kapsule control plane | Free | Free | | 1x DEV1-M node (3 vCPU, 4 GB) | EUR 14.50 | EUR 43.50 | | Object Storage (50 GB cached modules) | EUR 0.38 | EUR 1.14 | | Object Storage egress (500 GB/mo) | EUR 4.25 | EUR 12.75 | | Load Balancer (Envoy Gateway) | EUR 11.50 | EUR 34.50 | | DNS (GeoDNS) | — | EUR 1 | | **Total** | **EUR 31** | **EUR 93/mo** | For comparison: [JFrog Artifactory Cloud](https://jfrog.com/pricing/) starts at $150/month for a single region with 2 GB transfer. A self-hosted Artifactory on AWS with equivalent European coverage (3 regions, ELBs, EBS, data transfer) runs $400-600/month before you count engineering time. Even a single Athens instance on a t3.medium in eu-west-1 with an ALB costs $50-70/month for one region. EUR 93 buys three European regions with geo routing, automated TLS, GitOps deployment, and centralised observability. Sovereignty carries no cost premium. ## Decisions that earned their keep **Per-region caches, no replication.** Each region's Object Storage bucket warms independently from local usage. A module popular in Warsaw doesn't pre-populate in Paris. This eliminates cross-region data transfer costs, simplifies the GDPR story (no data moves between regions), and removes replication complexity. The trade-off — cold starts in under-used regions — is acceptable for a cache. A cache miss adds one round trip to upstream, then the module is cached for everyone in that region. **ALIAS record at the apex, not CNAME.** DNS zone apex records [cannot be CNAMEs](https://www.isc.org/blogs/cname-at-the-apex-of-a-zone/) (RFC 1034 section 3.6.2). Scaleway supports ALIAS records with `geo_ip` routing, giving geographic load distribution at the apex. The trade-off: Scaleway's health checks work on A/AAAA records but not ALIAS. A region going down continues receiving traffic from its geo IP match. The Go client's built-in fallback chain handles this — developers see a brief timeout, then fall through to `proxy.golang.org`. We chose simplicity over DNS-level failover. **Shared-bucket rate limiting, not per-IP.** Envoy Gateway applies 1000 requests per minute per proxy instance as a shared bucket. Per-IP rate limiting would require either logging client IPs (breaking our privacy architecture) or running an external Redis rate-limit service (adding cost and complexity). The shared bucket is less precise but maintains the zero-IP-logging guarantee. **DaemonSet proxy, not Deployment.** Envoy Gateway runs as a DaemonSet — one proxy pod per node. On a single-node cluster this makes no practical difference, but it guarantees the proxy scales with nodes if we ever scale up, without HPA configuration. ## GDPR enforcement at the architecture level Privacy compliance is not a policy document. It is a configuration choice. **Layer 1: Envoy Gateway access log format.** The format string omits `%DOWNSTREAM_REMOTE_ADDRESS%` and `%REQ(X-FORWARDED-FOR)%`. Client IPs never appear in access logs because the log format does not include them. ``` [%START_TIME%] "%REQ(:METHOD)% %REQ(X-ENVOY-ORIGINAL-PATH?:PATH)% %PROTOCOL%" %RESPONSE_CODE% %RESPONSE_FLAGS% %BYTES_RECEIVED% %BYTES_SENT% %DURATION% ``` **Layer 2: Alloy log pipeline.** Grafana Alloy's log processing stage regex-replaces any IPv4 or IPv6 address pattern with `[REDACTED]` before shipping to Loki. If an IP leaks through application logs, it is scrubbed before it reaches storage. The result: our observability stack (Mimir for metrics, Loki for logs, Grafana for dashboards) contains zero client IP addresses. Not because we delete them after collection, but because they never arrive. ## What we tried and rejected **Cross-region cache replication.** S3 Cross-Region Replication would pre-warm all caches from a single seed. We rejected it because: (1) it triples storage cost for marginal benefit — most modules are small and upstream latency is acceptable for cold fetches; (2) it creates cross-region data flows that complicate the data residency story; (3) it adds operational complexity for a problem that solves itself through usage patterns. **HTTP-01 ACME challenges.** Our initial setup used HTTP-01 for Let's Encrypt certificates. This requires the ACME challenge response to be routable through Envoy Gateway, which creates a circular dependency during initial deployment (no cert → no HTTPS → can't validate). We switched to DNS-01 challenges via a [Scaleway cert-manager webhook](https://github.com/scaleway/cert-manager-webhook-scaleway), which validates through DNS TXT records and works regardless of ingress state. **Prometheus per cluster.** A Prometheus server in each cluster would add 500 MB+ of memory per region. Grafana Alloy (DaemonSet, ~100 MB) scrapes all metrics locally and remote-writes to a centralised Mimir instance on [Heystaq](https://heystaq.com). One fewer stateful workload per cluster, one fewer thing to monitor. ## Operational reality **Deployment.** Push to `main`, Flux reconciles within 60 seconds. Per-region Kustomize overlays handle the differences (S3 bucket names, endpoints, hostnames). Adding a region is: copy a tofu environment, copy a Flux cluster entry, copy a Kustomize overlay, bootstrap Flux, add a geo IP DNS record. We did the third region (Warsaw) in under two hours. **Failover.** No automatic DNS failover — Scaleway's geo IP ALIAS doesn't support health checks. When a region goes down, the Go client fallback chain handles it: timeout on goproxy.eu, fall through to proxy.golang.org, builds continue. We've run deliberate failover tests. Client-side recovery takes 5-10 seconds per module fetch during the timeout, then builds proceed normally against Google's proxy. **Secrets.** SOPS with age encryption. Two age keys: one for our cluster secrets (API keys, credentials), one for Heystaq's cluster (observability config). Separate trust boundaries, both stored in Git encrypted. No external KMS dependency. **Monitoring.** Nine Grafana dashboards, 23 alert rules across 6 groups. Non-critical alerts route to Slack via GoAlert. Critical alerts (region down, certificate expiring, high error rate) escalate to SMS/voice. All metrics carry `cluster={fr-par,nl-ams,pl-waw}` labels for per-region filtering. ## The sovereignty part was free We did not design for the EU Cloud Sovereignty Framework. We designed for operability: GitOps, open-source components, declarative secrets, European hosting. When we [mapped the result against SEAL](https://gitlab.aknostic.com/aknostic/goproxy/-/blob/main/docs/seal-analysis.md), it scored well on all eight Sovereignty Objectives — not because we optimised for compliance, but because good engineering practices and sovereignty requirements converge. The practices that make a system auditable (everything in Git), portable (open-source components), and jurisdictionally contained (encrypted secrets, European infrastructure) are the same practices that make it operable. Sovereignty is a side effect of engineering discipline. ## The stack | Component | Version | License | Role | |-----------|---------|---------|------| | Athens | 0.15.x | MIT | Go module proxy | | Envoy Gateway | 1.3.0 | Apache-2.0 | Ingress, TLS, rate limiting | | Flux CD | 2.7.x | Apache-2.0 | GitOps | | cert-manager | 1.17.x | Apache-2.0 | TLS automation | | Redis | 7.x (Bitnami) | RSALv2 | SingleFlight dedup | | Grafana Alloy | 1.6.x | Apache-2.0 | Metrics/logs collection | | OpenTofu | 1.10.x | MPL-2.0 | Infrastructure as code | | Scaleway Kapsule | K8s 1.35 | — | Managed Kubernetes | No component is irreplaceable. Athens has alternatives (Goproxy, go-mod-proxy) though each has different configuration and storage backends. Envoy Gateway can be swapped for another Gateway API implementation. Flux for ArgoCD. Scaleway for any European cloud with managed Kubernetes and S3-compatible storage. Migration is never zero-effort, but nothing here creates a hard dependency. --- *goproxy.eu is a [Clouds of Europe](https://clouds-of-europe.eu) community project. [Source code](https://gitlab.aknostic.com/aknostic/goproxy) licensed under Apache-2.0.* --- ## Kubernetes Is Not American Infrastructure URL: https://clouds-of-europe.eu/content/policy-sovereignty/open-source-standards/kubernetes-is-not-american-infrastructure Author: Jurg van Vliet Published: 2026-02-27 Category: Policy & Sovereignty Type: Open Source & Standards Tags: kubernetes # The Origin Story Everyone Knows Google built Borg, its internal container orchestration system, in the early 2000s. A decade later, a team at Google distilled those ideas into an open-source project called Kubernetes, donated it to the newly formed Cloud Native Computing Foundation, and stepped back from sole ownership. That was 2014. In the eleven years since, Kubernetes has become the operating system of the cloud, and the story of its origin has become a kind of sovereignty anxiety for European organisations wondering whether their infrastructure depends on American goodwill. It does not. But understanding why requires looking past the creation myth and into how Kubernetes actually works as a project, a community, and an institution. # Governance That Outgrew Its Creator The CNCF is a project of the Linux Foundation. Its governing board includes representatives from companies across continents: European firms like SAP, Bosch, and Siemens sit alongside American, Chinese, and Japanese members. Technical decisions are made by Special Interest Groups whose membership is open and whose leadership is elected by contributors, not appointed by any single company. Google retains influence in this structure. Its engineers contribute heavily and its cloud platform benefits from Kubernetes adoption. But influence is not control. The Kubernetes release process, the API review process, and the security response process are all community-governed. No single company can block a release, reject a feature, or suppress a vulnerability disclosure. This governance model is not accidental. It was designed explicitly to prevent the scenario that European policymakers worry about: a single company holding critical infrastructure hostage. The CNCF's charter, its intellectual property framework, and its trademark policies all exist to make Kubernetes ungovernable by any one entity. That is the entire point. # The API Is the Sovereignty Layer Kubernetes' most important contribution to European digital sovereignty is not the software itself. It is the API. The Kubernetes API is a standard interface for describing computational workloads, storage, networking, and access control. Every conformant Kubernetes distribution implements this API. A deployment manifest that runs on Google's GKE runs identically on OVHcloud's managed Kubernetes, on Scaleway's Kapsule, on a self-managed cluster in a Frankfurt data centre, or on a k3s installation on a Raspberry Pi in someone's office. This is portability at the infrastructure layer, something the industry has promised and failed to deliver for decades. VMware promised it and then charged licensing fees that made migration economically irrational. OpenStack promised it and then fractured into incompatible distributions. Kubernetes actually delivered it, not because of altruism, but because the API conformance tests are rigorous and the ecosystem punishes non-compliance. For European organisations this means something concrete: choosing Kubernetes does not mean choosing Google. It means choosing an interface that European providers implement, European engineers operate, and European companies extend through operators and custom resources. The workloads are portable. The skills are portable. The operational tooling is portable. Moving from a US cloud provider to a European one is a migration project, not a rewrite. # European Kubernetes Is Already Real The narrative that Kubernetes is American infrastructure does not survive contact with the European cloud landscape. OVHcloud runs managed Kubernetes from data centres in France, Germany, Poland, and the UK. Scaleway operates Kapsule from Paris. IONOS provides Kubernetes from German infrastructure. Exoscale runs it from Switzerland. Hetzner, already popular for its pricing, offers Kubernetes from Finnish and German locations. These are not toy deployments or proofs of concept. They are production platforms serving European companies that have made deliberate sovereignty choices. Beyond managed offerings, European organisations run self-managed Kubernetes at significant scale. Public sector institutions, financial services companies, and healthcare providers operate clusters on European soil, managed by European teams, subject exclusively to European law. The software they run was contributed to by Google engineers, just as it was contributed to by engineers at Red Hat, VMware, Microsoft, independent contributors, and dozens of European companies whose names rarely appear in the origin story. # What "European" Software Actually Looks Like There is a persistent assumption in sovereignty discussions that European software must be written by Europeans, governed by Europeans, and hosted by Europeans to qualify as sovereign. This assumption is both impractical and philosophically incoherent. The internet protocols that European networks depend on were designed at American universities. The TLS encryption that protects European banking was standardised by the IETF, a US-incorporated organisation. The DNS system that resolves European domain names was invented by American researchers. None of this makes European internet infrastructure "American." Infrastructure becomes sovereign when the people operating it have full control over its configuration, its data, and its availability, regardless of who wrote the first line of code. Kubernetes fits this model precisely. European organisations deploying Kubernetes on European infrastructure, managed by European engineers, storing European data in European jurisdictions, are running sovereign infrastructure. The fact that the first commit came from a Google office in Mountain View is historically interesting and strategically irrelevant. # The Real Dependency to Watch If Kubernetes itself is not a sovereignty risk, what is? The answer is managed Kubernetes from hyperscalers. GKE, EKS, and AKS bundle Kubernetes with proprietary integrations: load balancers, identity systems, logging pipelines, and monitoring stacks that create quiet, accumulating lock-in. The Kubernetes API is portable but the ecosystem around a hyperscaler's managed offering is deliberately not. Moving off GKE means replacing Google's proprietary ingress controller, its workload identity federation, its operations suite, and dozens of small integrations that seemed convenient when they were adopted. This is where European sovereignty work should focus. Not on whether Kubernetes was created in America, but on whether the deployment is entangled with a specific American cloud provider's proprietary extensions. Building on the Kubernetes API with portable, open-source components, Gateway API instead of proprietary ingress, Prometheus instead of proprietary monitoring, Flux or ArgoCD instead of proprietary deployment pipelines, is what makes infrastructure genuinely sovereign. # Building on Solid Ground Kubernetes is the closest thing the industry has to neutral infrastructure. Its governance is multi-national, its API is a genuine standard, its implementations span continents and providers, and its license makes restriction impossible. Europe should engage with it not as a foreign dependency to be tolerated, but as shared infrastructure to be shaped. Contribute to the project. Run European Kubernetes distributions. Build operators that encode European operational practices. Participate in SIG governance. Maintain the European provider ecosystem that makes Kubernetes portable in practice, not just in theory. The foundation is solid. The question is not whether to build on it, but how ambitiously. #kubernetes #freedom-to-operate #european-cloud #open-source-governance #vendor-lock-in --- ## Does Go Belong to Google? URL: https://clouds-of-europe.eu/content/policy-sovereignty/open-source-standards/does-go-belong-to-google Author: Jurg van Vliet Published: 2026-02-26 Category: Policy & Sovereignty Type: Open Source & Standards **Does Go Belong to Google? The Question Europe Should Stop Asking** ## Origins and Ownership Go was created at Google in 2007 by Rob Pike, Robert Griesemer, and Ken Thompson. Google funded its development, hosts its module proxy, and employs many of its core contributors. By any conventional measure, Go is a Google project. But conventional measures miss something important about how programming languages actually work. When a language reaches a certain maturity, it stops belonging to its creator. Not legally, not through some transfer of intellectual property, but through the quiet, irreversible accumulation of community investment. Millions of lines of production code, thousands of contributors, hundreds of companies whose critical infrastructure depends on the language continuing to exist and evolve. Go crossed that threshold years ago. The question is not whether Google created it. The question is whether Google could meaningfully take it away, and the answer is no. ## What "Control" Actually Means Google controls the Go compiler, runs the module proxy, and employs the release team. These facts create a surface-level narrative of dependency that deserves more scrutiny than it typically receives. The Go specification is open and the compiler is BSD-licensed, which means anyone can fork it, build it, and distribute it without Google's permission or involvement. This is not theoretical. Alternative implementations exist: TinyGo compiles Go for microcontrollers and WebAssembly without using Google's compiler at all. The module proxy, proxy.golang.org, is a convenience rather than a requirement. Setting `GOPROXY=direct` tells the Go toolchain to fetch modules from their source repositories, bypassing Google entirely. European organisations can run Athens, an open-source module proxy, on their own infrastructure and after the initial cache warmup their builds never contact Google again. There is no technical reason why a European public Go proxy should not exist, and the Clouds of Europe community is exploring exactly that: a shared `go.clouds-of-europe.eu` that would give the continent its own module infrastructure, reducing a dependency to an architectural choice rather than a default. Compare this to genuinely controlled ecosystems. Oracle's Java required licensing negotiations for years. Microsoft's .NET was Windows-only for over a decade. Apple's Swift still cannot build Linux GUI applications without community effort that Apple does not support. Go's BSD license makes these kinds of restrictions structurally impossible. ## The European Open-Source Tradition Europe has a long and underappreciated tradition of shaping the tools it depends on without needing to have created them. Linux was created by a Finnish student and today it is the foundation of European public infrastructure, maintained by contributors from every continent, governed by a foundation incorporated in the US but operating as a genuinely global institution. No European policymaker seriously argues that Linux is "American software" because the Linux Foundation has a San Francisco address. The same pattern applies to PostgreSQL, born from a UC Berkeley research project, now the database of choice for European institutions precisely because no single company controls it. The PostgreSQL Global Development Group spans continents and European contributors are not guests in someone else's project. They are co-owners. Go follows this trajectory. European companies contribute to the compiler, maintain critical libraries, and run significant production workloads. The Go community in Europe is not consuming an American product. It is participating in a global commons. ## The Fork as Insurance Open-source licenses provide a guarantee that proprietary ecosystems cannot: the right to fork. This right is not symbolic, and it has been exercised repeatedly at moments that mattered. When Oracle acquired Sun Microsystems the MySQL community forked to MariaDB, and European Linux distributions switched within months. When HashiCorp changed Terraform's license from MPL to BSL the community forked to OpenTofu within weeks, and the project is now governed by the Linux Foundation with broad industry backing. Go's BSD license provides the same insurance. A hypothetical decision by Google to restrict Go would trigger a fork before the announcement finished circulating. The compiler is well-understood, the specification is public, and the community has the expertise to maintain an independent implementation indefinitely. This has never been necessary because Google has no incentive to restrict the language. Go's value to Google comes from its ecosystem: the libraries, the tools, the community of developers who choose it for production systems. Restricting the language would destroy the ecosystem that makes it valuable in the first place. ## What Europe Should Actually Worry About The sovereignty risk in programming languages is not about who created the compiler. It is about lock-in to proprietary ecosystems built on top of the language. Google Cloud's Go libraries tie applications to GCP. AWS SDK for Go ties applications to Amazon. These dependencies are real sovereignty concerns because they create migration friction by design. The language itself creates none. European organisations using Go should invest in standards-based libraries, portable abstractions, and infrastructure that runs anywhere. The Kubernetes ecosystem, built almost entirely in Go, demonstrates this principle clearly: the same operator binary runs on GKE, on Scaleway Kubernetes, on a bare-metal cluster in a Dutch data centre. The language does not care where it runs, and neither should the software built with it. ## Embracing the Commons Europe's path forward is not to build European programming languages. That would be an extraordinary misallocation of resources, reinventing a solved problem for reasons of flag-planting rather than engineering. The path forward is to participate vigorously in global open-source communities: to contribute, to maintain critical libraries, to run shared European infrastructure like module proxies and package mirrors, and to build expertise so deep that these communities cannot function without European involvement. This is how sovereignty works in open-source. Not through ownership, but through indispensability. Go does not belong to Google any more than Linux belongs to Linus Torvalds or the internet belongs to DARPA. Origins matter for history. Governance and participation matter for sovereignty. Europe should use Go confidently, contribute to it generously, and build production systems on it without apology. The language is a global commons, and Europe's place in that commons is earned through contribution, not guaranteed by geography. --- ## US Open-Source Software on European Terms URL: https://clouds-of-europe.eu/content/strategy-transition/strategic-planning/us-open-source-software-on-european-terms Author: Jurg van Vliet Published: 2026-02-26 Category: Strategy & Transition Type: Strategic Planning ## A Gift Worth Recognising Somewhere inside Target Corporation, one of the largest US retailers, a team of engineers built an on-call platform. They designed escalation policies, rotation schedules, alert routing. They solved the hard problems: who gets woken up at 3 AM, in what order, with what fallbacks. Then they released it all under the Apache 2.0 license. This deserves a moment of appreciation. A company with no obligation to share its internal tooling chose to give it away, irrevocably. Apache 2.0 is not "open-source until we change our minds." It is a permanent, unconditional grant. Target cannot revoke it, restrict European usage, or change the terms for existing code. Every commit to GoAlert is a gift that cannot be taken back. In an industry increasingly shaped by license rug-pulls, the bait-and-switch from open-source to "source available," and the quiet enclosure of community projects behind enterprise paywalls, releasing production-grade software under Apache 2.0 is an act of genuine generosity. This is what doing good looks like in software. European organisations building on GoAlert are not taking a risk. They are accepting a gift. ## Running It On Our Terms We run GoAlert on European infrastructure. The PostgreSQL database lives in our Kubernetes cluster. The GraphQL API serves from our network. User data, contact methods, escalation policies, rotation schedules never leave European jurisdiction. Target Corporation has no access to our instance, no telemetry connection, no control plane. The entire on-call topology is defined as Kubernetes Custom Resources, stored in Git, on our GitLab instance, applied through Flux. A declarative operator reconciles the desired state to GoAlert's API. The system of record is ours. The runtime is ours. The operational knowledge encoded in escalation policies and rotation schedules, the genuinely valuable part, lives entirely under European control. Self-hosted Apache 2.0 software is sovereignty in practice, not sovereignty in theory. ## The Twilio Question GoAlert sends notifications. For SMS and voice, it uses Twilio. This means a US company processes phone numbers and delivers messages through US-operated telecommunications infrastructure. Worth acknowledging. Also worth putting in perspective. An on-call system has two layers. The orchestration layer decides who to alert, when, through what escalation path. The delivery layer puts a message on a screen or makes a phone ring. The orchestration is the complex, valuable part. It encodes institutional knowledge about team structures, response priorities, and escalation logic built up over months. The delivery is a commodity. Twilio handles the commodity. It knows a phone number received a message. It does not know which service is down, what the incident looks like, or how your infrastructure is shaped. The blast radius of Twilio's US jurisdiction is limited to delivery metadata. Replacing Twilio with a European SMS gateway is straightforward engineering work. GoAlert supports multiple notification backends. The notification interface is a plugin point, not a load-bearing wall. Nobody has prioritised this yet because Twilio works reliably and the sovereignty exposure is narrow. That seems like a reasonable allocation of engineering attention. ## What Actually Threatens Sovereignty Contrast this with the mainstream alternative. PagerDuty, Opsgenie, and similar platforms run your entire on-call logic on US infrastructure. Your escalation policies, rotation schedules, alert history, incident timelines all live in a US data centre, operated by a US company, subject to the CLOUD Act. You cannot self-host. You cannot fork. You cannot migrate without rebuilding from scratch. The difference is structural. With GoAlert self-hosted, the US dependency is a replaceable delivery channel at the edge. With SaaS on-call platforms, the US dependency is the entire system. One architecture gives you a clear upgrade path. The other gives you vendor lock-in dressed up as convenience. European organisations should focus their sovereignty energy where the leverage is highest: owning the control plane, owning the data, owning the operational logic. A replaceable US delivery channel is a minor pragmatic concession compared to surrendering the whole stack to a SaaS provider that holds your data hostage. ## Walking the Path Accept Twilio for now. Not out of indifference, but because the architecture makes replacement easy when the time comes. The on-call logic runs in Europe. The data stays in Europe. The desired state lives in Git, versioned and auditable, on European infrastructure. The foundation is sound. The path to full sovereignty is short and well-marked. And it was made possible by engineers at a US retailer who chose to share their work with the world, no strings attached. That generosity is worth building on. --- ## Lock-In: A Room Without Doors URL: https://clouds-of-europe.eu/content/strategy-transition/business-cases/lock-in-a-room-without-doors Author: Jurg van Vliet Published: 2026-02-25 Category: Strategy & Transition Type: Business Cases Tags: europeancloud, lockin, portability, sovereignty, strategy ## You Don't Notice the Walls Going Up Nobody plans to get locked in. It happens gradually, one reasonable decision at a time. You pick a managed database because it saves operational effort. You adopt a proprietary queuing service because the integration is seamless. You sign a three-year reserved instance commitment because the discount is significant. You send your engineers to vendor certification courses because the training is subsidised. Each decision makes perfect sense in isolation. Together, they build a structure around you. By the time you realise you're enclosed, moving out is a project that nobody wants to fund. Cloud lock-in is often discussed as a single phenomenon, but it's actually four distinct forces. They reinforce each other, and understanding them separately is essential to managing them. ## Wall One: Technical Lock-In Technical lock-in is the most visible form. It happens when your systems depend on proprietary services that have no equivalent elsewhere. AWS Lambda, DynamoDB, Azure Cosmos DB, Google BigQuery. Each is a capable service. Each is available from exactly one provider. The mechanism is straightforward. You build your application around a proprietary API. Your code calls that API thousands of times. Your architecture assumes the service's specific behaviour, its latency characteristics, its consistency model, its scaling patterns. Replacing it means rewriting, not just reconfiguring. DynamoDB is a good example. It's a powerful NoSQL database, but its data model, query patterns, and capacity management are unlike anything else. Organisations that build heavily on DynamoDB report that migration requires not just moving data but redesigning how their applications think about storage. ScyllaDB, which offers a DynamoDB-compatible API specifically to ease migrations, describes the process as one of the most common pain points driving companies to seek alternatives. The counterpoint is Kubernetes. Because it's governed by the Cloud Native Computing Foundation and implemented by dozens of providers, a deployment manifest that works on Scaleway Kapsule works on OVHcloud, Hetzner, or AWS EKS. The same is true for PostgreSQL, Prometheus, and other foundation-governed projects. When the interface is a genuine standard with multiple implementations, technical lock-in largely disappears. The practical test is simple: could you run this component on a different provider without rewriting your application? If the answer is no, you've accepted technical lock-in. That might be a reasonable tradeoff, but it should be a conscious one. ## Wall Two: Data Lock-In Data lock-in is subtler than technical lock-in, and often more expensive to resolve. Your data is in a provider's region, stored in their format, integrated with their other services. Getting it out is technically possible but practically painful. The economics tell the story. AWS charges around €0.09 per gigabyte for data leaving their network. That sounds modest until you calculate what it means for a production database. Moving 10 terabytes of data costs roughly €900 in egress fees alone, before you account for the engineering time to validate the migration, maintain consistency during the transition, and verify nothing was lost. But egress fees are only the surface. The deeper problem is data gravity. Once your data lives in a provider's ecosystem, other services cluster around it. Your analytics pipeline reads from that database. Your machine learning models train on that data. Your backup systems, your audit logs, your compliance reports all reference it. Moving the data means moving or rebuilding everything that touches it. 37signals learned this firsthand. When David Heinemeier Hansson decided to leave AWS in 2023, the company was spending $3.2 million per year on cloud infrastructure. The migration to owned hardware ultimately saved them over $1.5 million annually. But the migration itself required months of careful work to extract data and rebuild the services that depended on AWS-specific storage and networking. Europe's regulators have noticed. The EU Data Act, which entered into application in September 2025, requires cloud providers to offer open interfaces and standard formats for data export. Starting January 2027, providers won't be allowed to charge switching fees at all. This is a recognition that data lock-in is a market failure that regulation needs to address. ## Wall Three: Contractual Lock-In Contractual lock-in works through commitments, discounts, and penalties. It's the least technical form of lock-in, but often the hardest to escape because it involves legal obligations rather than engineering challenges. The pattern is familiar. A cloud provider offers a 30 to 50 percent discount if you commit to one or three years of usage. The finance team loves the predictability. The procurement team appreciates the savings. Everyone signs. Now you're committed, and your negotiating position at renewal is weaker because migration during the contract term means paying for capacity you're not using. Broadcom's acquisition of VMware in 2023 showed how contractual lock-in can turn painful overnight. After the acquisition, Broadcom restructured VMware's licensing from perpetual licenses to mandatory subscriptions, bundled products customers didn't need, and enforced minimum 72-core requirements. The European Cloud Consortium reported price increases between 800 and 1,500 percent. Customers locked into VMware's ecosystem had limited options: pay the increase, or fund a migration to alternatives like Proxmox or Kubernetes while still paying VMware during the transition. Less dramatic but equally effective are the small contractual details. Termination fees. Data retention clauses that make it unclear who owns derivative data. Support tiers that bundle infrastructure management with the compute contract, so leaving the platform means losing operational support simultaneously. The antidote is shorter commitments and explicit exit terms. Every contract should answer: what happens when we want to leave? What format will our data be in? What are the costs? How long will the transition take? If the contract doesn't address these questions clearly, the omission is itself a form of lock-in. ## Wall Four: Skills Lock-In Skills lock-in is the least discussed form, but it's often the decisive one. When your team's expertise is entirely in one provider's ecosystem, switching providers means retraining everyone, and retraining is slow, expensive, and disruptive. Cloud providers understand this well. AWS has over a dozen certification programmes. Azure and Google Cloud have similar tracks. The certifications are genuinely valuable for working within those ecosystems. They're also specific to those ecosystems. An AWS Solutions Architect certification teaches you to think in AWS terms: regions, availability zones, VPCs, security groups, IAM policies. The concepts transfer partially, but the muscle memory, the tooling familiarity, and the troubleshooting instincts are vendor-specific. The compounding effect is significant. Teams that spend years building on one platform develop institutional knowledge that's hard to quantify and impossible to export. They know the quirks: which service has unexpected rate limits, which region has better availability, which support tier actually gets responses. This knowledge took years to accumulate and is worthless on another platform. Over time, hiring reinforces the pattern. Job descriptions ask for "3 years AWS experience." Candidates self-select. New hires bring more AWS knowledge, and the organisation's dependency deepens. Eventually, proposing a migration isn't just a technical discussion. It's a conversation about retraining an entire team, and nobody wants to have it. The alternative is investing in transferable skills. Kubernetes administration works across providers. PostgreSQL expertise applies everywhere PostgreSQL runs. Prometheus, Grafana, GitOps with Flux or Argo: these are skills that travel. Engineers who understand the underlying standards can work on any implementation. This doesn't mean ignoring provider-specific knowledge. It means building expertise in layers: standards at the foundation, provider-specific optimisations on top. When the foundation is portable, the provider-specific layer is thin enough to rebuild. ## How the Walls Reinforce Each Other What makes cloud lock-in so effective is that these four forces work together. Technical dependencies make data migration harder. Data gravity makes contractual negotiations weaker. Contractual commitments fund training in provider-specific skills. Provider-specific skills make technical dependencies feel natural rather than constraining. Breaking free from one wall while the others remain intact is exhausting. You might containerise your application to escape technical lock-in, but if your data is gravitationally bound and your team only knows one platform, the containerisation doesn't actually give you the freedom to move. This is why lock-in strategies need to address all four dimensions. Portable technology choices, standard data formats with tested export procedures, contracts with explicit exit terms, and teams whose core skills are transferable. ## The European Dimension For European organisations, cloud lock-in has an additional layer. When you're locked into a US hyperscaler, you're not just locked into a vendor. You're locked into a jurisdiction, a regulatory framework, and an economic relationship where value flows across the Atlantic. The EU Data Act's switching provisions, the ongoing discussions around EUCS (European Cybersecurity Certification Scheme), and the broader push for digital sovereignty all reflect a growing European consensus: lock-in isn't just a commercial inconvenience. It's a strategic vulnerability. Building on open standards, European providers, and transferable skills doesn't eliminate all risk. But it keeps the walls from going up around you. And that's what sovereignty means in practice: not isolation, but the maintained ability to choose. **Sources:** - [EU Data Act: Cloud Switching Requirements (Latham & Watkins, 2025)](https://www.lw.com/en/insights/eu-data-act-significant-new-switching-requirements-due-to-take-effect-for-data-processing-services) - [37signals Cloud Exit Savings (DHH)](https://world.hey.com/dhh/our-cloud-exit-savings-will-now-top-ten-million-over-five-years-c7d9b5bd) - [Broadcom VMware Price Increases 800-1,500% (The Register, 2025)](https://www.theregister.com/2025/05/22/euro_cloud_body_ecco_says_broadcom_licensing_unfair/) - [DynamoDB Migration Challenges (ScyllaDB)](https://www.scylladb.com/2024/03/06/dynamodb-how-to-move-out/) - [AWS Data Transfer Pricing](https://aws.amazon.com/ec2/pricing/on-demand/) #lockin #sovereignty #strategy #portability #europeancloud --- ## When Atlassian Goes Full SaaS-Monolith, Go Open Source URL: https://clouds-of-europe.eu/content/practice/implementation-patterns/when-atlassian-goes-full-saas-monolith-go-open-source Author: Jurg van Vliet Published: 2026-02-25 Category: Practice Type: Implementation Patterns ## The OpsGenie Problem Atlassian is retiring OpsGenie as a standalone product, absorbing it into the Jira Service Management monolith. For European infrastructure teams, that means buying into an ever-growing bundle of non-EU SaaS services just to keep your on-call schedules running. We chose not to follow. ## GoAlert We are replacing OpsGenie with GoAlert, an open-source on-call management system originally built by Target Corporation. It does what OpsGenie does: on-call scheduling, escalation policies, SMS and voice notifications with two-way acknowledgment. It's a single Go binary backed by PostgreSQL. We run it on Scaleway in France, next to our Grafana stack. ``` Grafana Alert Rules → Grafana Alertmanager → GoAlert Webhook → Notifications ↑ SMS ack (1a/1c) ``` Grafana fires alerts, GoAlert routes them to whoever is on call, Twilio delivers SMS and voice. Engineers acknowledge by replying `1a` to the SMS. When Grafana resolves the alert, GoAlert auto-closes it. What convinced us was the simplicity. GoAlert has a clean GraphQL API, handles deduplication by alert summary, and supports generic webhook ingestion. Any system that can send an HTTP POST can create alerts. We use this for forwarding Datadog and CloudWatch alerts from customers who haven't migrated their monitoring yet. ## GitOps for Incident Management GoAlert's weakness is operational management. Services, escalation policies, schedules, and user notification rules are configured through a web UI or direct API calls. For a team that manages everything through Git, that wasn't going to work. So we built `goalert-provisioning`, a Kubernetes operator that reconciles GoAlert configuration from Custom Resource Definitions: ```yaml apiVersion: goalert.heystaq.com/v1alpha1 kind: GoAlertService metadata: name: atlas-critical namespace: org-atlas spec: serviceName: "Atlas CRITICAL" escalationPolicyRef: name: aknostic-critical namespace: aknostic integrationKeys: - name: datadog-critical secretRef: name: atlas-integration-keys key: datadog-critical ``` You push to Git, Flux syncs to the cluster, the operator creates the service in GoAlert and writes the integration key token to a Kubernetes Secret. Onboarding a new tenant takes a pull request. We onboarded five tenants this way, each with critical and non-critical alert routing, in under an hour per tenant. The operator covers the full lifecycle: services and integration keys, escalation policies, on-call schedules, rotations, user accounts and contact methods, even system-level admin configuration. This mirrors the pattern we already use for Grafana resources through the Grafana Operator. The entire incident management stack, from metric collection to alert routing to on-call notification, lives in version control. ## gctl: a CLI for Model-Based RCA The operator solved provisioning, but we also wanted to bring large language models into incident response. That required a programmatic interface to GoAlert, so we built `gctl`. ```bash gctl oncall gctl alert list gctl query '{ services { nodes { name } } }' ``` It's a thin GraphQL client for checking who's on call, listing active alerts, or querying services. But the reason we built it is that we expose our GoAlert to our agents. An RCA investigation usually starts at the end, with the alert(s). It needs to find patterns in the incident response management, just as important as the other components of our observability stack. The model gets access to the alert context, related metrics in Mimir, logs in Loki, and traces in Tempo. It correlates signals across these sources and proposes a root cause. It doesn't replace the engineer's judgment, but it handles the tedious part: cross-referencing dashboards, crafting log queries, chasing through trace spans. The model operates within the same tenant boundaries enforced by the platform, querying Mimir and Loki with the correct `X-Scope-OrgID` header. ## What We Gained Replacing OpsGenie wasn't just about avoiding a forced SaaS migration. It changed how we work. Alert rules, escalation policies, on-call schedules, contact points all live in one repository with one review process and one audit log. When someone asks why an alert went to a particular person, the answer is a Git commit. Each customer gets a Kubernetes namespace containing their GoAlert CRDs, Grafana CRDs, and SOPS-encrypted secrets. Onboarding is a directory copy with find-and-replace. GoAlert is Apache 2.0 licensed, the operator and gctl are open source, and the underlying infrastructure runs on standard Kubernetes with S3-compatible storage. Alert data, on-call schedules, notification logs are all stored in PostgreSQL on Scaleway France. The one remaining US dependency is Twilio for SMS and voice delivery, which we might replace, but there are more important workloads to repatriate first. The goalert-provisioning operator and gctl are available at [gitlab.aknostic.com/aknostic/goalert-provisioning](https://gitlab.aknostic.com/aknostic/goalert-provisioning). The operator installs via Helm and works with any GoAlert instance. gctl installs with `go install`. If you're running GoAlert and want GitOps provisioning, or if you're looking at alternatives to OpsGenie, we'd welcome contributors and feedback. [GoAlert Provisioning Operator](https://gitlab.aknostic.com/aknostic/goalert-provisioning) [GoAlert][(https://goalert.me/) *This is a follow-up to [Building European Observability: Multi-Location Synthetic Monitoring with GitOps](https://clouds-of-europe.eu/content/practice/implementation-patterns/building-european-observability-multi-location-synthetic-monitoring-with-gitops), documenting the next phase of the Heystaq platform.* --- ## SEAL Assessment: Clouds of Europe URL: https://clouds-of-europe.eu/content/policy-sovereignty/policy-regulation/seal-assessment-clouds-of-europe Author: Jurg van Vliet Published: 2026-02-16 Category: Policy & Sovereignty Type: Policy & Regulation The EU Cloud Sovereignty Framework (published Oct 2025) grades across 8 Sovereignty Objectives (SOV-1 to SOV-8), each scored SEAL-0 to SEAL-4. Here's where this project lands: ## Overall Score: SEAL-2 (Data Sovereignty) EU law applies and data stays in the EU, but material non-EU dependencies remain. ## SOV-1: Strategic Sovereignty (15%) — SEAL-3 | Factor | Status | |---|---| | Ownership | Private project, no non-EU investors | | Governance | Self-hosted GitLab on EU infrastructure | | License | EUPL-1.2 (specifically European) | | Capital | No dependency on non-EU capital | Strong. European license, EU-hosted source control, no foreign governance exposure. ## SOV-2: Legal & Jurisdictional (10%) — SEAL-2 | Factor | Status | |---|---| | Infrastructure jurisdiction | French law (Scaleway) | | OAuth providers | Google, GitHub, LinkedIn — all US, subject to CLOUD Act | | CI/CD | Self-hosted GitLab (EU) | The three OAuth providers create exposure to US jurisdiction. User tokens and profile data flow through US services. Email magic links are EU-only, but social login is the primary path. ## SOV-3: Data & AI (10%) — SEAL-3 | Factor | Status | |---|---| | Database | PostgreSQL on Scaleway, fr-par region | | Backups | Scaleway S3, fr-par, SOPS-encrypted | | Data residency | All data in France | | AI/ML | None used | Data never leaves the EU. Encryption at rest (SOPS/AGE) and in transit (TLS). No AI/ML processing, so no data sovereignty concerns there. ## SOV-4: Operational (15%) — SEAL-3 | Factor | Status | |---|---| | Infrastructure management | OpenTofu, Flux GitOps — all EU-hosted | | Monitoring | HeyStaq Grafana (EU) | | Support staff | EU-based | | Autonomous operation | Can operate without non-EU dependencies | Full operational control. GitOps model means no external operator access. Monitoring is EU-based. Could operate independently if needed. ## SOV-5: Supply Chain (20%) — SEAL-1 | Factor | Status | |---|---| | Base container images | node:22-alpine from Docker Hub (US) | | PostgreSQL image | ghcr.io/cloudnative-pg/postgresql (GitHub, US) | | npm packages | 95%+ US-maintained | | Hardware | Scaleway (French), but underlying chips are non-EU | | Container registry | Scaleway (EU) for built images | Weakest area. Every build pulls base images from US registries. The npm ecosystem is overwhelmingly US-based. This is the industry-wide problem — no European project can score high here without significant investment in mirroring infrastructure. ## SOV-6: Technology (15%) — SEAL-3 | Factor | Status | |---|---| | Open source stack | 100% (Next.js, PostgreSQL, Kubernetes, Flux) | | Vendor lock-in | None — standard K8s, portable across providers | | Proprietary dependencies | Zero | | Open APIs/protocols | HTTPS, SQL, SMTP — all open standards | Excellent technology sovereignty. Entire stack is open source, runs on standard Kubernetes, and could be migrated to any EU cloud provider. ## SOV-7: Security & Compliance (10%) — SEAL-3 | Factor | Status | |---|---| | GDPR | Privacy-by-design, no tracking, consent-based | | Encryption | SOPS/AGE (at rest), TLS/Let's Encrypt (in transit) | | Secrets management | SOPS-encrypted, Flux auto-decrypts | | Network security | NetworkPolicies deployed | | Rate limiting | Distributed via Memcached | | Security scanning | TruffleHog, SOPS validation in CI | Good compliance posture. Recent improvements (NetworkPolicies, rate limiting) strengthen this. Missing: container scanning (Trivy/Snyk) and SAST. ## SOV-8: Environmental (5%) — SEAL-1 | Factor | Status | |---|---| | Green energy | No documentation | | Scaleway DC-5 | Adiabatic cooling, PUE ~1.3 | | Carbon reporting | None | No explicit sustainability commitments or documentation. Scaleway's French DCs are relatively efficient but this isn't documented or leveraged. ## Weighted Score Breakdown | SOV | Weight | Score | Weighted | |---|---|---|---| | SOV-1 Strategic | 15% | 3 | 0.45 | | SOV-2 Legal | 10% | 2 | 0.20 | | SOV-3 Data | 10% | 3 | 0.30 | | SOV-4 Operational | 15% | 3 | 0.45 | | SOV-5 Supply Chain | 20% | 1 | 0.20 | | SOV-6 Technology | 15% | 3 | 0.45 | | SOV-7 Security | 10% | 3 | 0.30 | | SOV-8 Environmental | 5% | 1 | 0.05 | | **Total** | **100%** | | **2.40 / 4.00 (60%)** | ## Top 3 Improvements for SEAL-3 1. **Supply chain mirroring** (SOV-5, 20% weight) — Mirror base images to Scaleway registry, add SBOM generation. Moves from SEAL-1 to SEAL-2. 2. **EU authentication** (SOV-2, 10% weight) — Add Keycloak or EU-based IdP alongside existing OAuth. Moves from SEAL-2 to SEAL-3. 3. **Sustainability documentation** (SOV-8, 5% weight) — Document Scaleway's energy efficiency, add to README. Low effort, moves from SEAL-1 to SEAL-2. --- **Sources:** - [SUSE Cloud Sovereignty Framework Self Assessment](https://www.suse.com/cloud-sovereignty-framework-assessment/) - [EU Cloud Sovereignty Framework - European Commission](https://commission.europa.eu/document/download/09579818-64a6-4dd5-9577-446ab6219113_en) - [EU's Cloud Sovereignty SEAL Ranking - InfoQ](https://www.infoq.com/news/2025/11/eu-seal-framework-governance/) - [Safespring SEAL Self-Assessment](https://www.safespring.com/blogg/2025/2025-11-the-eu-just-defined-sovereign-cloud-here-is-our-score/) - [From Policy to Proof - SUSE](https://www.suse.com/c/from-policy-to-proof-cracking-the-black-box-of-digital-sovereignty/) --- ## Writing Runbooks: Our Alert-Driven Documentation Approach URL: https://clouds-of-europe.eu/content/practice/getting-started-guides/writing-runbooks-our-alert-driven-documentation-approach Author: Jurg van Vliet Published: 2026-02-09 Category: Practice Type: Getting Started Guides Documentation has a reputation problem. Teams write comprehensive operational guides during calm periods, then never look at them again. When something actually breaks at 2am, the runbook is outdated, doesn't cover the actual failure mode, or requires fifteen minutes of reading before you find the relevant command. Our runbooks are different. Each one exists because an alert fired and someone had to figure out what to do. They're not aspirational documentation about what *might* happen—they're battle-tested procedures for what *did* happen. This article shares our approach: the template we use, how alerts link directly to runbooks, and the culture that keeps runbooks useful rather than decorative. ## The Core Insight: Runbooks Are Executable Procedures The difference between documentation and a runbook is the verb tense. Documentation describes: "The PostgreSQL cluster uses CloudNativePG for high availability. Replication is configured synchronously. Failover happens automatically when the primary becomes unavailable." A runbook instructs: "Check cluster status: `kubectl get cluster postgres-cluster -n clouds-of-europe`. If replicas show 0, check pod events for scheduling failures. If PVC is stuck in zone-locked state, delete the PVC and pod—CNPG will recreate them." Documentation tells you how things work. Runbooks tell you what to do when things don't work. This distinction matters at 2am when your phone is buzzing and your brain is operating at 60% capacity. You don't need architecture explanations. You need commands to copy-paste, decisions to make, and escalation paths when you're stuck. ## The Template Every runbook follows the same structure: ```markdown # Runbook: AlertName ## Alert Details | Alert | Severity | Pending | Threshold | |-------|----------|---------|-----------| | AlertName | critical | 5m | Description of trigger condition | **Data Source:** Where the metric comes from ## Quick Diagnosis 1. First thing to check (with link to dashboard) 2. Second thing to check (with command) 3. Third thing to check (with log query) ## Common Causes ### 1. Most Likely Cause **Symptoms:** What you'll observe **Check:** Command or query to confirm **Fix:** Exact steps to resolve ### 2. Second Most Likely Cause ... ## Emergency Procedures What to do if things are critically broken and you need to restore service immediately, even with temporary fixes. ## Verification How to confirm the issue is actually resolved. ## Escalation When to page someone else, and who. ## Related Alerts Other alerts that often fire together or share root causes. ``` This structure is deliberate: **Alert Details** at the top lets responders confirm they're looking at the right runbook. The threshold reminds them what triggered the alert. **Quick Diagnosis** gives three fast checks to understand the situation. These should take under two minutes total. **Common Causes** lists the actual reasons this alert has fired historically, ordered by frequency. The first cause listed is what you should check first. **Emergency Procedures** provides the "break glass" options when you need to restore service immediately, even imperfectly. **Verification** prevents premature closure. The alert might stop firing, but is the underlying issue actually fixed? **Related Alerts** helps responders understand cascading failures. If PostgreSQL is down, expect application errors too. ## Real Examples ### PostgreSQL No Replica This runbook was written after we woke up to find our production database running without a replica—HA completely degraded without anyone noticing until the alert fired. ```markdown ## Alert Details | Alert | Severity | Pending | Threshold | |-------|----------|---------|-----------| | PostgreSQL No Replica | critical | 5m | Streaming replicas < expected | ## Quick Diagnosis 1. **Check PostgreSQL Cluster Status** ```bash kubectl get cluster postgres-cluster -n clouds-of-europe -o yaml | grep -A 20 "status:" ``` 2. **Check Pod Status** ```bash kubectl get pods -n clouds-of-europe -l cnpg.io/cluster=postgres-cluster -o wide ``` 3. **Check Pod Events** ```bash kubectl get events -n clouds-of-europe --field-selector involvedObject.name=postgres-cluster-2 ``` ## Common Causes ### 1. Replica Pod Pending - PVC Zone Affinity **Symptoms:** Pod pending, PVC bound to node in unavailable zone **Resolution:** 1. Delete the stuck PVC: `kubectl delete pvc postgres-cluster-2 -n clouds-of-europe` 2. Delete the pending pod: `kubectl delete pod postgres-cluster-2 -n clouds-of-europe` 3. CNPG will recreate both in an available zone ``` The PVC zone affinity issue was the actual cause when this alert first fired. Scaleway's block storage is zone-locked—a PVC created in `fr-par-1` can only attach to nodes in `fr-par-1`. When autoscaling removed that node, the replica couldn't reschedule. The fix is counterintuitive (delete the PVC!) but correct. Without this runbook, responders would spend thirty minutes trying to figure out why the pod won't schedule. ### App Heap Memory High This runbook came from a memory leak that crept in after a dependency update. We watched heap usage climb over hours until the alert fired. ```markdown ## Quick Diagnosis 1. **Check Application Dashboard** - Look at "Heap Usage" panel in Node.js Runtime section - Check "Process Memory" for RSS trends 2. **Check Pod Resource Usage** ```bash kubectl top pods -n clouds-of-europe ``` 3. **Check Application Logs for Memory Warnings** ```logql {namespace="clouds-of-europe"} |~ "(?i)(memory|heap|oom)" ``` ## Common Causes ### 1. Memory Leak **Symptoms:** Steady increase in heap over time, never decreasing **Fix:** - Rollback recent deployment if leak started after deploy - Restart pods as temporary mitigation: ```bash kubectl rollout restart deployment/clouds-of-europe-app -n clouds-of-europe ``` ### 2. High Traffic / Load **Symptoms:** Heap spikes correlate with request rate **Fix:** - Scale up replicas to distribute load: ```bash kubectl scale deployment/clouds-of-europe-app -n clouds-of-europe --replicas=3 ``` ``` The "restart pods as temporary mitigation" note is important. It's not a real fix—the leak is still there. But at 2am, restoring service matters more than root cause analysis. The runbook acknowledges this pragmatism while making clear it's temporary. ### Flux Reconciliation Failing This runbook is longer because Flux failures have many possible causes. GitOps is powerful, but when the reconciliation loop breaks, deployments stop. ```markdown ## Common Causes ### 1. Git Repository Access Issues **Symptoms:** GitRepository source not ready **Check:** ```bash flux get sources git -A kubectl describe gitrepository flux-system -n flux-system ``` ### 2. Invalid Manifests / Helm Values **Symptoms:** Kustomization or HelmRelease fails to apply **Check:** ```bash flux logs --level=error kubectl get events -n flux-system --sort-by='.lastTimestamp' ``` **Fix:** - Review recent Git commits for syntax errors - Validate manifests locally: ```bash kustomize build gitops/infrastructure/production ``` ## Emergency Procedures ### If Production Deployment Blocked 1. **Check if issue is blocking production changes:** ```bash flux get kustomization clouds-of-europe-app -n flux-system ``` 2. **If urgent, apply manually (temporary):** ```bash # Only for emergencies - GitOps will reconcile later kubectl apply -f ``` 3. **Document manual intervention** for later cleanup ``` The "apply manually" instruction is deliberately marked as emergency-only. It breaks the GitOps principle, but sometimes you need to ship a hotfix *now*. The runbook permits this while emphasizing it's an exception. ## Linking Alerts to Runbooks Every Grafana alert includes a `runbook_url` annotation: ```yaml - uid: production-postgres-no-replica title: "Production PostgreSQL No Replica" annotations: description: "Production PostgreSQL cluster has fewer streaming replicas than expected." runbook_url: "https://github.com/aknostic/clouds-of-europe/blob/main/docs/runbooks/postgres-no-replica.md" summary: "Production PostgreSQL replica missing" ``` When the alert fires, the runbook URL appears in the notification. In Grafana's alert UI, it's a clickable link. In Slack or SMS notifications, it's visible text. This removes the "where's the runbook?" friction. The alert tells you something is wrong *and* where to find help. You're one click from the exact procedure you need. We use full GitHub URLs rather than relative paths because notifications go to multiple channels—Slack, SMS, email. An absolute URL works everywhere. ## The Alert-Runbook Lifecycle Here's how runbooks actually get written: 1. **Alert fires for the first time.** Someone responds, figures out what's wrong, fixes it. 2. **After the incident**, that person writes a runbook. Not during—after. The immediate priority is restoration. Documentation happens when you're calm. 3. **The runbook goes in Git**, in `docs/runbooks/`, named after the alert. 4. **The alert configuration gets updated** to include `runbook_url` pointing to the new runbook. 5. **Next time the alert fires**, the responder follows the runbook. If it's incomplete or wrong, they update it. This is the key: runbooks evolve through use. The first version captures what worked once. Subsequent incidents refine it. After three or four firings, the runbook covers the common cases well. Runbooks written in advance, before any incident, are speculation. They might be helpful. They might be wrong. You won't know until something breaks. Our approach ensures every runbook has been tested in production conditions at least once. ## The Threshold Update Story A small commit tells a bigger story: ``` commit a720e3d9 Update heap memory alert threshold to 90% in runbook ``` The alert threshold was changed from 80% to 90% in Grafana. The runbook still said 80%. Someone noticed during an incident—the runbook said the alert fires at 80%, but it had actually fired at 90%. This kind of drift is inevitable. Alert thresholds get tuned. Runbooks get stale. The fix is simple: when you notice a discrepancy, fix it immediately. The commit took thirty seconds. The underlying lesson: runbooks are code. They live in Git. They get reviewed and merged like any other change. When the system changes, the runbook changes too. ## Synthetic Monitoring Test Procedure One runbook is unusual: it documents how to *intentionally* trigger alerts. ```markdown ## Testing Production Alerts To test synthetic monitoring alerts, block the probe IPs using an Envoy Gateway SecurityPolicy: ```bash # Suspend Flux first to prevent auto-reconciliation flux suspend helmrelease clouds-of-europe -n flux-system # Block the 3 probe IPs cat <99.5%" but provides no guarantees ([source](https://learn.microsoft.com/en-us/azure/aks/free-standard-pricing-tiers)) - Premium tier ($0.60/hr) required for Long-Term Support - Hidden costs in Azure CNI networking **GCP GKE:** - Free tier credit effectively covers one cluster - Autopilot mode adds ~20% overhead vs self-managed nodes - Egress pricing 20% higher than AWS for first TB ([source](https://holori.com/egress-costs-comparison/)) ### STACKIT Assessment **Potential advantages:** - German/Austrian data centers only — **true** EU data sovereignty - Backed by Schwarz Group (Lidl/Kaufland) — €7B+ annual revenue, serious investment - Free egress traffic currently ([source](https://www.stackit.de/en/pricing/cloud-services/)) - Potentially 15-20% cheaper than Scaleway for compute **Concerns:** - **Maturity**: Relatively new entrant (public since 2022) - **Ecosystem lock-in**: Proprietary tooling, less community support - **Limited regions**: Germany + Austria only - **Documentation**: Sparse compared to hyperscalers - **Kubernetes version**: Check if they support latest K8s releases --- ## Data Sovereignty Reality Check | Provider | HQ | US Law Exposure | True Sovereignty | |----------|-----|-----------------|------------------| | **Scaleway** | France | None | ✅ Yes | | **STACKIT** | Germany | None | ✅ Yes | | **AWS EU Sovereign** | USA | CLOUD Act applies¹ | ❌ Illusory | | **Azure/GCP EU** | USA | CLOUD Act applies | ❌ No | ¹ Despite AWS's "European Sovereign Cloud" marketing, legal experts confirm US jurisdiction cannot be contractually waived. French MP Philippe Latombe: *"AWS cloud cannot be sovereign because it is subject to the US FISA and Cloud Act"* ([source](https://eliatra.com/blog/the-sovereignty-illusion-why-awss-european-cloud-cannot-escape-us/)) --- ## Sources - [Scaleway Containers Pricing](https://www.scaleway.com/en/pricing/containers/) - [Scaleway Virtual Instances Pricing](https://www.scaleway.com/en/pricing/virtual-instances/) - [AWS EKS Pricing](https://aws.amazon.com/eks/pricing/) - [AWS EKS Control Plane Pricing Analysis](https://clustercost.com/blog/aws-eks-control-plane-pricing-2025/) - [Azure AKS Pricing Tiers](https://learn.microsoft.com/en-us/azure/aks/free-standard-pricing-tiers) - [GKE Pricing Guide](https://www.devzero.io/blog/gke-pricing) - [STACKIT Kubernetes Engine](https://www.stackit.de/en/product/kubernetes/) - [STACKIT Pricing](https://www.stackit.de/en/pricing/cloud-services/) - [Cloud Egress Cost Comparison](https://holori.com/egress-costs-comparison/) - [AWS Sovereignty Illusion Analysis](https://eliatra.com/blog/the-sovereignty-illusion-why-awss-european-cloud-cannot-escape-us/) - [EU Cloud Sovereignty Risks](https://unit8.com/resources/eu-cloud-sovereignty-emerging-geopolitical-risks/) - [European Cloud Pricing Analysis](https://european.cloud/2025/06/a-basic-look-at-pricing-of-european-cloud-vendors/) --- ## Managed Observability in Europe: Why We Chose HeyStaq Over Self-Hosted Prometheus URL: https://clouds-of-europe.eu/content/strategy-transition/collaborative-infrastructure/managed-observability-in-europe-why-we-chose-heystaq-over-self-hosted-prometheus Author: Jurg van Vliet Published: 2026-02-04 Category: Strategy & Transition Type: Collaborative Infrastructure We ran self-hosted Prometheus, Loki, and Grafana for three months. Then we deleted 38,000 lines of configuration and migrated to a managed service. This article isn't a vendor pitch—it's an honest accounting of what self-hosted observability actually costs, and why a European managed platform made sense for our stage. The commit that removed our self-hosted stack deleted 84 files across Mimir, Loki, Grafana, Tempo, blackbox-exporter, and their supporting infrastructure. Each of those files represented decisions, debugging sessions, and operational knowledge accumulated over weeks. We walked away from all of it. Here's why. ## The Self-Hosted Dream The appeal of self-hosted observability is obvious, especially for a European sovereignty platform: - **Complete control**: You own your metrics, your retention policies, your access controls - **No vendor lock-in**: Standard protocols (Prometheus remote_write, Loki push API) mean you can move - **Cost predictability**: Compute and storage costs, not per-metric pricing - **Sovereignty by default**: Data never leaves your infrastructure We built a comprehensive stack on our management cluster: - **Mimir** for long-term metrics storage (Prometheus-compatible) - **Loki** for log aggregation - **Grafana** for visualization and alerting - **Tempo** for distributed tracing - **Prometheus Agent** on each workload cluster, shipping via remote_write - **Promtail** for log shipping - **Blackbox Exporter** for synthetic monitoring Each component had its own Helm chart, ConfigMaps for configuration, Secrets for credentials, ServiceMonitors for self-monitoring, and dashboards for visibility. The GitOps repository grew accordingly. ## The Hidden Costs What the architecture diagrams don't show is the operational tax. ### Memory Tuning Mimir and Loki are memory-hungry applications. Not in the "allocate 4GB and forget it" way, but in the "tune your ingester batch sizes, query concurrency, and compaction schedules or watch OOM kills" way. Our Mimir ingesters would periodically get OOM-killed during compaction. The solution wasn't more memory—it was understanding the interaction between: - `-ingester.max-global-series-per-user` - `-blocks-storage.bucket-store.max-chunk-pool-bytes` - `-server.grpc-max-recv-msg-size-bytes` - Pod memory limits Getting this right took multiple iterations, each requiring a deployment, observation period, and adjustment. ### Storage Math Prometheus metrics aren't small. A typical Kubernetes cluster generates thousands of time series. Each series has samples at your scrape interval (usually 15-30 seconds). Retention compounds quickly. Our calculations: - ~50,000 active series across both clusters - 15-second scrape interval - 8 bytes per sample (timestamp + value) - 30-day retention That's roughly 70GB of raw data, before accounting for indexes and metadata. Mimir's block storage format is efficient, but you're still looking at significant S3 costs and query latency for historical data. We spent more time on retention policies than we'd like to admit. Which metrics warrant 30 days? Which can be downsampled? Which can be dropped after a week? These questions don't have universal answers—they depend on your debugging patterns and compliance requirements. ### The Upgrade Treadmill Observability components release frequently. Mimir, Loki, and Grafana each have their own release cadence, their own breaking changes, their own deprecation timelines. Staying current isn't optional. Security patches, performance improvements, and bug fixes matter for infrastructure you depend on. But each upgrade requires: 1. Reading release notes for breaking changes 2. Testing in a non-production environment (if you have one) 3. Planning the upgrade sequence (some components have ordering dependencies) 4. Executing the upgrade 5. Verifying everything still works 6. Rolling back if it doesn't This isn't a one-time cost. It's a recurring tax on engineering time, every few weeks. ### Configuration Complexity Our prometheus-agent configuration alone was 300 lines. Metric relabeling rules, cAdvisor filtering, ServiceMonitor selectors, resource limits. Each line represented a decision and potential failure mode. When metrics stopped appearing in Grafana, the debugging path included: 1. Is the target being scraped? (Check Prometheus targets page) 2. Are ServiceMonitor labels matching? (Check label selectors) 3. Is remote_write working? (Check Prometheus logs) 4. Is Mimir accepting the samples? (Check Mimir logs) 5. Is the series being dropped by a relabeling rule? (Check relabel configs) 6. Is the series exceeding cardinality limits? (Check Mimir limits) 7. Is the query correct? (Check PromQL syntax) Multiply this by every component in the stack. ## The Migration Decision After three months, we had a working stack. It ingested metrics, stored logs, fired alerts. But we also had: - A significant portion of our GitOps repository dedicated to observability - Regular time spent on observability maintenance instead of product work - Occasional outages where our monitoring was down and we couldn't monitor our monitoring The irony of not knowing your observability stack is broken because your observability stack is broken is not lost on anyone who's lived it. We evaluated our options: **Keep self-hosting**: Accept the operational cost as part of sovereignty **Use a US-based managed service**: Datadog, Grafana Cloud US regions **Use a European managed service**: Find a provider with EU data residency The third option hadn't been obvious initially. The observability market is dominated by US companies. But we found HeyStaq, a multi-tenant observability platform built by Aknostic, a European company running European infrastructure. ## What HeyStaq Provides HeyStaq is built on the same open-source components we were running: - **Mimir** for metrics (Prometheus-compatible) - **Loki** for logs - **Grafana** for visualization and alerting - **Multi-location synthetic monitoring** from European cities The difference: Aknostic operates it. They handle the OOM tuning, the storage scaling, the upgrades, the high availability. We ship metrics and logs via standard protocols. **Endpoints:** | Service | URL | |---------|-----| | Mimir (metrics) | `https://mimir.heystaq.com/api/v1/push` | | Loki (logs) | `https://loki.heystaq.com/loki/api/v1/push` | **Multi-tenancy:** The `X-Scope-OrgID: clouds-of-europe` header isolates our data from other tenants. We can't see their metrics; they can't see ours. **Authentication:** mTLS client certificates for metric and log shipping. Google OAuth for Grafana access. ## The Migration Migrating was straightforward because we were already using standard protocols. Prometheus Agent's `remote_write` and Promtail's Loki client just needed new endpoints and credentials. ### Prometheus Agent Configuration ```yaml prometheus: prometheusSpec: remoteWrite: - url: https://mimir.heystaq.com/api/v1/push headers: X-Scope-OrgID: clouds-of-europe writeRelabelConfigs: - targetLabel: environment replacement: production - targetLabel: kubernetes_cluster replacement: production queueConfig: capacity: 10000 maxShards: 10 maxSamplesPerSend: 5000 sampleAgeLimit: 5m tlsConfig: cert: secret: name: heystaq-client-mtls key: tls.crt keySecret: name: heystaq-client-mtls key: tls.key ``` The `sampleAgeLimit: 5m` was learned through experience—after a prometheus-agent restart, it would try to re-send buffered samples, some of which were older than Mimir's out-of-order window. Dropping samples older than 5 minutes prevents rejection errors. ### mTLS Setup HeyStaq uses mutual TLS for authentication. We receive a client certificate and key, store them as a Kubernetes Secret (SOPS-encrypted in Git), and reference them in the prometheus-agent and promtail configurations. ```yaml apiVersion: v1 kind: Secret metadata: name: heystaq-client-mtls namespace: monitoring type: kubernetes.io/tls stringData: tls.crt: | -----BEGIN CERTIFICATE----- -----END CERTIFICATE----- tls.key: | -----BEGIN RSA PRIVATE KEY----- -----END RSA PRIVATE KEY----- ``` One gotcha: HeyStaq's servers use a public CA (Let's Encrypt). The CA certificate they provide is for server-side client verification, not for verifying the server. Including it in the client's CA bundle causes TLS failures. We learned this through debugging, not documentation. ### Deleting the Old Stack With shipping confirmed working, we deleted the self-hosted stack: ``` 84 files changed, 38,373 deletions(-) ``` Mimir configuration: gone. Loki configuration: gone. Grafana deployment: gone. Tempo: gone. Postgres cluster for Grafana state: gone. S3 credentials for metric storage: gone. Alert rules: migrated to HeyStaq's Grafana. The GitOps repository became dramatically simpler. ## What We Gained ### Operational Simplicity We no longer debug OOM kills in our observability stack. We no longer plan Mimir upgrades. We no longer calculate retention storage. Someone else does that—someone whose job is operating observability infrastructure, not building a community platform. ### Reliability HeyStaq runs redundant infrastructure across multiple availability zones. Our self-hosted stack ran on a single management cluster. When that cluster had issues, our monitoring had issues. Now, our monitoring is independent of our workload infrastructure. When our production cluster has problems, we can still see metrics and logs to diagnose them. ### Synthetic Monitoring HeyStaq runs Blackbox Exporter instances in Paris, Amsterdam, and Warsaw. They probe our endpoints every 30 seconds from three European vantage points. This is genuinely useful for detecting regional issues—and it's infrastructure we don't operate. ### Grafana Operator CRDs We converted our dashboards and alerts to Grafana Operator Custom Resource Definitions. These live in Git, are deployed via Kustomize, and provision directly into HeyStaq's Grafana instance: ```yaml apiVersion: grafana.integreatly.org/v1beta1 kind: GrafanaAlertRuleGroup metadata: name: synthetic-monitoring namespace: org-clouds-of-europe spec: instanceSelector: matchLabels: org: clouds-of-europe folderRef: synthetic-monitoring interval: 1m rules: - uid: af9j72o3pay2ob title: "EndpointDown" condition: C for: 1m labels: severity: "critical" ``` This gives us GitOps for observability configuration while using managed infrastructure for execution. The CRDs are portable—if we eventually return to self-hosting, the same files work. ## What We Lost ### Complete Control We can't modify Mimir's ingestion limits. We can't add custom recording rules at the Mimir level. We can't change Loki's retention beyond what the platform offers. These constraints haven't mattered yet, but they could. ### Cost Transparency Self-hosted costs were predictable: compute, storage, network. Managed platform pricing is less transparent—we pay based on our plan, and scaling is a conversation rather than a Terraform variable. ### Debugging Depth When something's wrong with metric ingestion on a self-hosted stack, you can read Mimir's logs, check its metrics, examine its configuration. With a managed platform, you open a support ticket. For a small team, this is fine. For larger organizations with dedicated observability teams, it might be limiting. ## The Sovereignty Calculation For a European digital sovereignty platform, the managed vs. self-hosted question has an additional dimension: where does data go? With HeyStaq: - Metrics and logs are stored on European infrastructure - The company operating the infrastructure is European - Data never leaves EU jurisdiction - GDPR applies to the provider relationship This is meaningfully different from using Grafana Cloud's EU region (still a US company, subject to US jurisdiction) or Datadog (same concern). We're still dependent on a third party. That's a trade-off against pure sovereignty. But the alternative—operating all infrastructure ourselves—has its own risks: reliability issues, security gaps from slow patching, and opportunity cost from infrastructure work instead of product work. ## When to Self-Host Our decision isn't universal. Self-hosted observability makes sense when: **You have a platform team**: Dedicated engineers who can own the operational burden, respond to issues, and stay current on upgrades. **You have strict compliance requirements**: Some regulations require not just data residency but operational control. Managed platforms may not satisfy auditors. **You need custom capabilities**: Recording rules, exotic metric types, integration with internal systems that don't play well with multi-tenant platforms. **Cost scales better**: At very high metric volumes, self-hosted can be cheaper than managed per-metric pricing. But the crossover point is higher than most expect once you account for engineering time. For an early-stage European platform with a small team, none of these applied. The managed path was clearly better. ## The Hybrid Model We didn't fully abandon self-hosted. Our architecture is: - **Shipping**: Self-managed Prometheus Agent and Promtail in each cluster - **Storage and query**: Managed HeyStaq (Mimir, Loki) - **Visualization and alerting**: Managed Grafana with GitOps-provisioned configuration - **Synthetic monitoring**: Managed blackbox-exporter probes This hybrid gives us control over what gets shipped (our metric relabeling rules, our ServiceMonitors) while offloading the stateful, complex parts (storage, query, high availability). The Prometheus Agent's `remote_write` is a portability guarantee. If we outgrow HeyStaq or want to return to self-hosting, we change an endpoint URL. The shipping configuration stays the same. ## Key Takeaways **Self-hosted observability is operationally expensive.** Memory tuning, storage calculation, upgrade management, configuration debugging—these costs are real and recurring. **European managed observability exists.** You don't have to choose between sovereignty and operational simplicity. Look for European providers with EU infrastructure. **Standard protocols enable migration.** Prometheus remote_write and Loki push API mean you're not locked in. Your shipping configuration works with self-hosted or managed backends. **GitOps works with managed platforms.** Grafana Operator CRDs let you version-control dashboards and alerts while using managed Grafana. You get the benefits of infrastructure-as-code without operating the infrastructure. **Know your stage.** Early-stage teams should focus on product, not observability infrastructure. As you grow, the calculus may change. Build with portability in mind. --- *This article documents work done on the Clouds of Europe platform in January 2026. --- ## Right-Sizing Kubernetes for European Startups: Our 2-Node HA Production Setup URL: https://clouds-of-europe.eu/content/practice/implementation-patterns/right-sizing-kubernetes-for-european-startups-our-2-node-ha-production-setup Author: Jurg van Vliet Published: 2026-02-04 Category: Practice Type: Implementation Patterns The Kubernetes community has a scaling problem—not the technical kind, but a cultural one. Conference talks showcase clusters with hundreds of nodes. Best practices assume you have a platform team. Resource examples default to generous allocations that work great in enterprise budgets. But what about the European startup running a Next.js application with a PostgreSQL database and a caching layer? What about the team that needs production reliability but can't justify €500/month in compute costs for an early-stage product? We run Clouds of Europe on a 2-node Scaleway Kapsule cluster. This article explains how we made that work: the QoS strategy that prevents OOM kills while maximizing node utilization, the anti-affinity rules that spread pods for availability, and the cost calculation that justified scaling down from three nodes to two. ## The Cost Reality Let's start with numbers. Scaleway's PRO2-XXS instances (2 vCPU, 8GB RAM) cost approximately €20/month each. Our initial production setup used three nodes across three availability zones: - 3 nodes × €20 = €60/month for compute - Plus storage, networking, load balancer When we analyzed actual resource usage, we found we were running at 30-40% CPU utilization on average. The third node existed for theoretical resilience, not practical necessity. Scaling to two nodes saved roughly €40/month (accounting for the full infrastructure delta). That's not transformative money, but for an early-stage project, €480/year funds other things—domain renewals, email services, the occasional debugging tool subscription. More importantly, the exercise forced us to think rigorously about resource allocation. The real savings came from the efficiency improvements, not just the node reduction. ## Understanding Kubernetes QoS Classes Kubernetes assigns every pod a Quality of Service class based on its resource configuration. This class determines what happens when nodes run low on resources: **Guaranteed**: Requests equal limits for both CPU and memory. These pods are never killed due to resource pressure (unless they exceed their own limits). They get exactly what they asked for, no more, no less. **Burstable**: Requests are lower than limits, or only some resources have limits. These pods can use more than their requests if capacity is available, but they're candidates for eviction when nodes are under pressure. **BestEffort**: No resource requests or limits set. These pods get whatever's left over and are first to be killed when resources are scarce. The conventional wisdom is "set requests equal to limits for production workloads." This gives you Guaranteed QoS and predictable behavior. But it also means you're reserving resources you might not use. ## Our Strategy: Guaranteed Memory, Burstable CPU Here's the insight that changed our approach: memory and CPU behave differently under pressure. **Memory is binary.** When a process needs memory and can't get it, bad things happen—OOM kills, data corruption, undefined behavior. There's no graceful degradation. If your application needs 512MB during a traffic spike, it needs 512MB. **CPU is throttleable.** When a process needs more CPU than available, it slows down. Requests take longer. But nothing crashes. The kernel scheduler ensures every process gets its fair share based on requests, and anything above that is best-effort. This means the optimal configuration for cost-conscious production is: ```yaml resources: requests: cpu: "100m" # Low request for scheduling memory: "512Mi" # Full memory requirement limits: cpu: "1000m" # Allow bursting to 10x request memory: "512Mi" # Same as request = Guaranteed for memory ``` This gives us: - **Burstable QoS overall** (because CPU requests < limits) - **Guaranteed memory behavior** (because memory requests = limits) - **Efficient scheduling** (because CPU requests are low) - **Headroom for spikes** (because CPU limits are high) ## Applying This Across the Stack We applied this pattern to every workload in our cluster. ### Application (Next.js) ```yaml resources: requests: cpu: "100m" # Low request for scheduling, can burst to limit memory: "512Mi" # Increased for Prisma operations limits: cpu: "1000m" # Allow bursting for API requests and Prisma operations memory: "512Mi" ``` The Next.js application is bursty by nature. Page renders and API calls spike CPU briefly, then go idle. Setting CPU requests at 100m (0.1 cores) means we're only "reserving" 10% of a core for scheduling purposes. But when a request comes in that needs computation—a complex Prisma query, server-side rendering, image processing—it can burst up to a full core. Memory is set at 512Mi for both requests and limits. Prisma's connection pool and Next.js's server-side caching need predictable memory. We can't afford OOM kills during traffic spikes. ### PostgreSQL (CloudNativePG) ```yaml resources: requests: cpu: 100m # Low request for scheduling, can burst to limit memory: 512Mi limits: cpu: 500m # Allow bursting to 500m when needed memory: 512Mi ``` PostgreSQL mostly sits idle waiting for queries. When queries arrive, they can be CPU-intensive (joins, aggregations, index scans). The 100m request means PostgreSQL doesn't block scheduling, while the 500m limit provides headroom for complex operations. Memory is critical for PostgreSQL—shared buffers, work memory, connection state. We set requests = limits at 512Mi. The database will never be evicted due to memory pressure. ### Memcached and Mcrouter ```yaml # Memcached resources: requests: memory: "128Mi" cpu: "20m" limits: memory: "128Mi" cpu: "100m" # Mcrouter resources: requests: memory: "32Mi" cpu: "30m" limits: memory: "32Mi" cpu: "100m" ``` Cache workloads are memory-bound. Memcached's entire purpose is holding data in RAM. CPU usage is minimal—parsing requests, serialization. We give them guaranteed memory with burstable CPU. ## The Scheduling Math Why does this matter for scheduling? Kubernetes schedules pods based on requests, not limits. When you ask for 500m CPU, the scheduler reserves 500m on a node—even if your actual usage is 50m. On a 2-vCPU node (2000m total), here's the difference: **Guaranteed approach (requests = limits):** - Application: 500m CPU - PostgreSQL: 500m CPU - Memcached: 100m CPU - Mcrouter: 100m CPU - System overhead: ~200m - **Total reserved: 1400m** - **Remaining for scheduling: 600m** **Our approach (low CPU requests):** - Application: 100m CPU - PostgreSQL: 100m CPU - Memcached: 20m CPU - Mcrouter: 30m CPU - System overhead: ~200m - **Total reserved: 450m** - **Remaining for scheduling: 1550m** The second approach leaves 2.5x more CPU available for additional pods or for cluster autoscaler decisions. When actual usage spikes, workloads burst into the remaining capacity. When usage is low (most of the time), we're not paying for idle reservation. ## Pod Anti-Affinity for High Availability Running on two nodes creates an obvious concern: what happens if one node fails? We use pod anti-affinity to spread replicas: ```yaml affinity: podAntiAffinity: preferredDuringSchedulingIgnoredDuringExecution: - weight: 100 podAffinityTerm: labelSelector: matchLabels: app: clouds-of-europe-app topologyKey: kubernetes.io/hostname ``` This tells Kubernetes: "prefer to schedule application pods on different nodes." If we run 2 replicas, they'll land on separate nodes. If one node fails, one replica survives. We use `preferredDuringSchedulingIgnoredDuringExecution` rather than `required` because we want flexibility. If both replicas must run and only one node is healthy, we'd rather have 2 replicas on one node than 0 replicas waiting for a second node. For PostgreSQL, CloudNativePG handles this automatically: ```yaml affinity: enablePodAntiAffinity: true topologyKey: "topology.kubernetes.io/zone" podAntiAffinityType: "preferred" ``` The primary and replica are scheduled in different availability zones when possible. If the zone with the primary fails, the replica promotes automatically. ## CloudNativePG: Database HA Without the Complexity Speaking of PostgreSQL, CloudNativePG deserves mention. It's a Kubernetes operator that manages PostgreSQL clusters with: - Automatic failover (primary dies, replica promotes) - Continuous WAL archiving to S3 - Point-in-time recovery - Zero-downtime updates via switchover Our configuration: ```yaml spec: instances: 2 # Primary + sync replica primaryUpdateStrategy: unsupervised primaryUpdateMethod: switchover backup: retentionPolicy: "30d" barmanObjectStore: destinationPath: "s3://coe-prod-pgbackups/backups" s3Credentials: accessKeyId: name: postgres-s3-credentials key: ACCESS_KEY_ID secretAccessKey: name: postgres-s3-credentials key: SECRET_ACCESS_KEY ``` With synchronous replication enabled, every committed transaction exists on two instances before acknowledgment. If the primary fails, we lose zero transactions. WAL archiving to Scaleway S3 means we can recover to any point in the last 30 days. This is genuine production resilience—not "we have backups somewhere" but "we can recover to any second of the last month with zero data loss." ## When Not to Do This This approach has limits. Here's when you should use larger clusters and Guaranteed QoS: **Latency-sensitive workloads**: If your application has strict P99 latency requirements, CPU throttling during contention is unacceptable. Set CPU requests = limits. **Multi-tenant clusters**: If untrusted workloads share your cluster, burstable pods can be starved by noisy neighbors. Guaranteed QoS provides isolation. **Regulated environments**: Some compliance frameworks require dedicated resources. Check your requirements. **Databases with heavy write loads**: Our PostgreSQL handles modest traffic. If you're doing thousands of writes per second, don't skimp on CPU. The pattern works for early-stage products with variable traffic, where cost efficiency matters and occasional CPU throttling is acceptable. ## Monitoring the Trade-offs We watch several metrics to ensure our lean configuration isn't causing problems: **CPU throttling**: `container_cpu_cfs_throttled_seconds_total` shows when containers hit their CPU limits. Some throttling is expected; sustained throttling means limits are too low. **Memory pressure**: `container_memory_working_set_bytes` vs limits. If working set approaches limits consistently, increase memory before OOM kills happen. **Node pressure**: `kube_node_status_condition{condition="MemoryPressure"}` and `DiskPressure`. If nodes are under pressure, workloads risk eviction. **Pod evictions**: `kube_pod_status_reason{reason="Evicted"}`. Any eviction of our workloads indicates the configuration is too aggressive. In practice, we've seen zero evictions and minimal CPU throttling. The bursty nature of web applications means contention is rare—traffic spikes don't perfectly align across all pods. ## The European Context Why does this matter specifically for European startups? European digital sovereignty isn't just about data location; it's about building technology that serves European values. Efficient, sustainable, appropriately-sized infrastructure fits that ethos better than over-provisioned clusters burning electricity. ## Key Takeaways **Memory and CPU are different.** Set memory requests = limits for guaranteed memory. Set CPU requests low with higher limits for efficient scheduling with burst capacity. **Understand QoS classes.** Guaranteed isn't always better. Burstable with guaranteed memory gives you predictable stability with scheduling efficiency. **Use pod anti-affinity.** On small clusters, spreading pods across nodes is essential for availability. Make it preferred, not required, for scheduling flexibility. **CloudNativePG is production-ready.** You can run PostgreSQL on Kubernetes with proper HA, automated failover, and point-in-time recovery. S3 backups make it affordable. **Monitor the trade-offs.** Watch throttling, memory pressure, and evictions. The configuration that works at low traffic might need adjustment as you grow. **Right-size for your reality.** The Kubernetes community's defaults assume enterprise scale. Question every resource configuration against your actual needs. --- *This article documents work done on the Clouds of Europe platform in January 2026. --- ## From GitHub to GitLab: A Practical Guide to CI/CD Independence URL: https://clouds-of-europe.eu/content/practice/tools-templates/from-github-to-gitlab-a-practical-guide-to-ci-cd-independence Author: Jurg van Vliet Published: 2026-02-03 Category: Practice Type: Tools & Templates There's something awkward about building a European digital sovereignty platform on Microsoft-owned infrastructure. GitHub Actions is convenient—deeply integrated, well-documented, generous free tier—but every workflow run happens on Azure machines in regions Microsoft chooses. Your source code, build logs, and deployment secrets all flow through American infrastructure. For Clouds of Europe, this contradiction is our raison d'être. We migrated our entire CI/CD pipeline to a self-hosted GitLab instance, and this article shares the practical lessons: the security scanning setup, the Docker build detour through Kaniko and back, and the workflow patterns that made selective deployments straightforward. ## Why GitLab, Not Just "Not GitHub" The obvious question: why not Gitea, Forgejo, or another lightweight Git server? The answer is CI/CD maturity. GitLab's pipeline system has evolved for over a decade. The runner ecosystem is stable. The documentation is comprehensive. And critically, you can self-host it on infrastructure you control. We run our GitLab instance on European infrastructure managed by Aknostic. The runners execute on European servers. Build artifacts stay in European storage. There's no ambiguity about jurisdiction—GDPR applies, and American subpoenas don't. This isn't paranoia. For organizations handling European citizen data, regulatory clarity matters. "Our CI/CD runs on Microsoft Azure" is a compliance conversation you don't want to have with a German regulator. ## The Pipeline Architecture Our pipeline has three stages: ```yaml stages: - security - build - deploy ``` Each stage is in its own file under `.gitlab/ci/`, included from the main `.gitlab-ci.yml`. This modularity makes it easy to understand what each stage does and to modify stages independently. The main configuration also defines pipeline inputs—dropdown menus in GitLab's "Run Pipeline" UI that let you skip stages or select deployment targets: ```yaml spec: inputs: start_with: description: "Stage to start from (skip earlier stages)" default: "security" options: - "security" - "build" - "deploy" deploy_env: description: "Environment to deploy" default: "test" options: - "test" - "production" ``` This seemingly minor feature turned out to be valuable. When you need to re-deploy production without rebuilding (maybe Flux got stuck, maybe you're testing deployment scripts), you can select `start_with: deploy` and `deploy_env: production`. No waiting for security scans and Docker builds you don't need. ## Security Scanning: TruffleHog and SOPS Validation GitHub Actions has a marketplace of security scanning actions. GitLab has... less of that. But it turns out that building your own security stage isn't hard, and the result is more transparent. Our security stage runs three jobs: ### TruffleHog Secret Scanning ```yaml trufflehog-scan: stage: security image: trufflesecurity/trufflehog:latest variables: GIT_DEPTH: 0 # Full clone for history scanning script: - | # Determine scan range based on pipeline type if [ "$CI_PIPELINE_SOURCE" = "merge_request_event" ]; then BASE_SHA="$CI_MERGE_REQUEST_TARGET_BRANCH_SHA" echo "Scanning MR changes from $BASE_SHA to HEAD" elif [ -n "$CI_COMMIT_BEFORE_SHA" ] && [ "$CI_COMMIT_BEFORE_SHA" != "0000000000000000000000000000000000000000" ]; then BASE_SHA="$CI_COMMIT_BEFORE_SHA" echo "Scanning push from $BASE_SHA to HEAD" else BASE_SHA="HEAD~10" echo "Scanning last 10 commits" fi trufflehog git file://. --since-commit="$BASE_SHA" --only-verified --fail --no-update ``` The key insight is `--only-verified`. TruffleHog can detect patterns that look like secrets (API keys, tokens, passwords), but many of these are false positives—example configurations, test fixtures, documentation. The `--only-verified` flag tells TruffleHog to actually test credentials against their services (AWS, GitHub, Slack, etc.) and only fail if they're real and active. This makes the scan dramatically more useful. You're not wading through hundreds of "this looks like it might be a secret" warnings. You're getting actionable "this is a real credential that works right now" alerts. ### Hardcoded Secrets Check TruffleHog catches committed credentials, but it doesn't catch everything. We have a simpler check that looks for patterns specific to our codebase: ```yaml hardcoded-secrets-check: stage: security image: alpine:latest script: - | FOUND_ISSUES=0 # Check for plain text appPassword in config files if grep -r 'appPassword:.*"[a-zA-Z0-9]' --include="*.yaml" --include="*.yml" . 2>/dev/null | \ grep -v -E "(template|values|example|gitops|node_modules|\.git|tests|docs|helm)" | grep -q .; then echo "Found plain text appPassword" FOUND_ISSUES=1 fi # Check for .plain.yaml files if find . -name "*.plain.yaml" ! -path "./.git/*" 2>/dev/null | grep -q .; then echo "Found plain text secret files (*.plain.yaml)" FOUND_ISSUES=1 fi if [ $FOUND_ISSUES -eq 0 ]; then echo "No hardcoded secrets detected" else exit 1 fi ``` This catches our specific anti-patterns: YAML files with `appPassword` that aren't in the expected locations (templates, examples), and `.plain.yaml` files that someone forgot to encrypt before committing. ### SOPS Validation The third check ensures our encrypted secrets are actually encrypted: ```yaml sops-validation: stage: security image: alpine:latest script: - | # Find YAML files in sops-secrets directories ENCRYPTED_FILES=$(find gitops/ -path "*/sops-secrets/*.yaml" ! -name "kustomization.yaml" 2>/dev/null || true) for file in $ENCRYPTED_FILES; do if ! grep -q "ENC\[AES" "$file" && ! grep -q "sops:" "$file"; then echo "File $file in sops-secrets/ doesn't appear to be SOPS encrypted" exit 1 fi echo "$file is properly encrypted" done ``` The `! -name "kustomization.yaml"` exclusion is important—we learned this the hard way. Kustomization files in `sops-secrets/` directories are configuration, not secrets. They shouldn't be encrypted, and checking them was causing false positives. ## The Kaniko Detour Now we get to the interesting part: Docker builds. GitHub Actions runners have Docker available by default. GitLab runners... it's complicated. The standard approach is Docker-in-Docker (DinD): run a Docker daemon as a sidecar service, and connect to it from your build job. This works, but it requires privileged runners—the Docker daemon needs capabilities that aren't available in unprivileged containers. Privileged runners are a security concern. A malicious job could potentially escape the container and access the host. For a self-hosted GitLab instance, this means trusting every job that runs on your infrastructure. We tried Kaniko, Google's tool for building container images without Docker: ```yaml build-and-push: stage: build image: name: gcr.io/kaniko-project/executor:v1.23.2-debug entrypoint: [""] before_script: - | # Create Kaniko config for registry authentication mkdir -p /kaniko/.docker echo "{\"auths\":{\"${REGISTRY}\":{\"auth\":\"$(echo -n nologin:${SCW_SECRET_KEY} | base64)\"}}}" > /kaniko/.docker/config.json script: - | /kaniko/executor \ --context="${CI_PROJECT_DIR}" \ --dockerfile="${CI_PROJECT_DIR}/Dockerfile" \ --destination="${IMAGE_BASE}:${CI_COMMIT_REF_SLUG}-${SHORT_SHA}" \ --cache=true \ --cache-repo="${IMAGE_BASE}/cache" ``` Kaniko works differently: it unpacks the base image, executes Dockerfile instructions as file operations, and repacks the result. No Docker daemon, no privileged mode. The problem? Kaniko's authentication handling is awkward. The config file needs to be in `/kaniko/.docker/`, which is read-only in the executor image. We tried mounting it from the project directory with `--kaniko-dir`, but that introduced other permissions issues. Registry authentication that worked fine with `docker login` required careful base64 encoding with Kaniko. After spending a day fighting Kaniko's quirks, we stepped back and asked: what's the actual risk we're mitigating? Our GitLab runners are dedicated to our projects. We control what jobs run on them. The privileged mode concern is real for shared runners where untrusted code might execute, but that's not our situation. We reverted to Docker-in-Docker: ```yaml build-and-push: stage: build image: docker:24 services: - docker:24-dind variables: DOCKER_HOST: tcp://docker:2375 DOCKER_TLS_CERTDIR: "" DOCKER_BUILDKIT: "1" before_script: - docker info - echo "$SCW_SECRET_KEY" | docker login $REGISTRY -u nologin --password-stdin ``` The key configuration is `DOCKER_HOST: tcp://docker:2375` and `DOCKER_TLS_CERTDIR: ""`. This disables TLS between the build job and the DinD sidecar. In a trusted network (which a Kubernetes pod's localhost is), TLS adds complexity without meaningful security benefit. The lesson: sometimes the straightforward solution is the right one. Kaniko solves a real problem—building images in environments where you can't run privileged containers—but if you control your runners, the added complexity isn't worth it. ## Deployment: OpenTofu and Flux The deploy stage handles infrastructure provisioning (OpenTofu) and application deployment (Flux GitOps). Here's the test environment deployment: ```yaml deploy-test: stage: deploy image: alpine:latest environment: name: test url: https://test.clouds-of-europe.eu script: - | # Initialize OpenTofu with Scaleway S3 backend cd infrastructure/opentofu/environments/test tofu init -upgrade \ -backend-config="bucket=coe-opentofu-state" \ -backend-config="key=test/terraform.tfstate" \ -backend-config="region=fr-par" \ -backend-config="endpoint=https://s3.fr-par.scw.cloud" \ -backend-config="access_key=${SCW_ACCESS_KEY}" \ -backend-config="secret_key=${SCW_SECRET_KEY}" tofu plan -out=tfplan tofu apply -auto-approve tfplan # Extract kubeconfig tofu output -raw kubeconfig > /tmp/kubeconfig.yaml export KUBECONFIG=/tmp/kubeconfig.yaml # Bootstrap or reconcile Flux if kubectl get deployment source-controller -n flux-system &>/dev/null; then flux reconcile source git flux-system -n flux-system else flux bootstrap gitlab \ --hostname=gitlab.aknostic.com \ --owner=clouds-of-europe \ --repository=clouds-of-europe \ --branch=main \ --path=gitops/clusters/test \ --token-auth fi artifacts: paths: - kubeconfig-test.yaml expire_in: 7 days ``` A few things worth noting: **OpenTofu over Terraform**: OpenTofu is the open-source fork of Terraform created after HashiCorp's license change. For a sovereignty-focused platform, using truly open-source infrastructure tooling aligns with our values. **Kubeconfig as artifact**: After deployment, we save the kubeconfig file as a pipeline artifact. This means you can download cluster access credentials directly from the GitLab UI—useful for debugging, and the 7-day expiration means credentials don't persist forever. **Flux idempotence**: The script checks whether Flux is already installed before bootstrapping. This makes the deployment job idempotent—you can run it multiple times without breaking the cluster. **Environment-specific secrets**: The SOPS Age key differs between test and production (`SOPS_AGE_KEY_TEST` vs `SOPS_AGE_KEY_PRODUCTION`). Each environment can only decrypt its own secrets. ## Production Deployment: Automatic When You Want It Production deployment requires explicit opt-in, but we made it automatic when you've already decided: ```yaml deploy-production: rules: # Run automatically when DEPLOY_ENV=production - if: $CI_PIPELINE_SOURCE == "web" && $DEPLOY_ENV == "production" # Manual trigger for other web runs - if: $CI_PIPELINE_SOURCE == "web" when: manual ``` If you run a pipeline from the web UI and select `deploy_env: production`, the deployment starts immediately—no clicking through manual gates. But if you run a normal pipeline without that explicit selection, production deployment requires a manual click. This balances safety with efficiency. You don't accidentally deploy to production, but when you intend to, you're not clicking through unnecessary confirmation dialogs. ## What We Gained Beyond independence, the migration gave us tangible improvements: **Clearer stage separation**: Security scanning, builds, and deployments are distinct. A build failure doesn't require re-running security scans. A deployment re-run doesn't require rebuilding. **Better artifact management**: Docker images push to Scaleway's registry in Paris. State files live in Scaleway S3. Kubeconfigs save as pipeline artifacts. Everything has a clear location. **Workflow flexibility**: The `start_with` and `deploy_env` inputs let operators skip stages and target environments without editing pipeline code. **Transparent security**: Instead of trusting a marketplace action, we see exactly what TruffleHog and our custom checks do. When something fails, the debug path is clear. ## What We Lost Honesty requires acknowledging the trade-offs: **GitHub Actions' ecosystem**: The marketplace has thousands of actions. GitLab's ecosystem is smaller. We wrote more shell scripts. **Documentation and community**: Stack Overflow has more GitHub Actions answers. When something breaks, you're more likely to find someone who's seen it before. **Integration convenience**: GitHub Actions workflows can reference other repositories, use GitHub's secrets management, trigger on GitHub events. We had to build some of this ourselves. For Clouds of Europe, independence outweighed these inconveniences. For projects without regulatory requirements or philosophical commitments to European infrastructure, GitHub Actions might be the better choice. The point isn't that GitLab is universally superior—it's that migration is achievable when independence matters. ## Key Takeaways **Self-hosted CI/CD is achievable.** GitLab's runner model is mature. The pipeline syntax is well-documented. You can migrate from GitHub Actions without heroic effort. **Build your own security scanning.** TruffleHog with `--only-verified` is more useful than marketplace actions that generate noise. Custom checks for your codebase's specific anti-patterns catch what generic tools miss. **Question your assumptions about privileged runners.** Kaniko solves a real problem, but if you control your runners, Docker-in-Docker with proper configuration might be simpler. **Use pipeline inputs for workflow flexibility.** The ability to skip stages and select deployment targets without editing code makes operations smoother. **Save kubeconfigs as artifacts.** Having cluster access available from the pipeline UI is valuable for debugging and emergency access. --- *This article documents work done on the Clouds of Europe platform in January 2026.* --- ## The Next Cloud Revolution: Why Kubernetes Operators Change Everything URL: https://clouds-of-europe.eu/content/practice/getting-started-guides/the-next-cloud-revolution-why-kubernetes-operators-change-everything Author: Jurg van Vliet Published: 2026-02-02 Category: Practice Type: Getting Started Guides Amazon and Microsoft didn't win the cloud war with better virtual machines. They won with developer convenience and a fundamental shift from capex to opex. Click a button, get a database. Call an API, send an email. No hardware procurement, no capacity planning, no upfront investment. The infrastructure disappeared behind managed services that simply worked, billed by the hour. European cloud companies saw this and tried to copy it. Build another AWS. Build another Azure. Build another GCP. But they lacked the commitment—the grit and the financial backing—to compete in a race defined by the hyperscalers. Billions in R&D. Thousands of engineers. Decades of runway. The European challengers brought knives to a gunfight and wondered why they kept losing. Here's the thing: we don't have to fight that battle anymore. We can redefine the game. Rewrite the rules entirely. ## The New Distribution Model Kubernetes changed everything, but not in the way most people think. The real revolution wasn't containers or orchestration—it was the operator pattern. Operators encode operational knowledge into software. They watch, reconcile, and heal. They turn complex stateful systems into something that behaves like a managed service. CloudNativePG gives you production-grade PostgreSQL. Strimzi gives you Kafka. Redis, RabbitMQ, MongoDB—operators exist for all of them. These aren't toys. They're battle-tested systems running critical workloads, maintained by communities and vendors with deep expertise. This is the paradigm shift: **software distribution is becoming the new managed service**. Instead of paying a hyperscaler for their proprietary database service, you deploy an operator that manages the database for you. The operational complexity doesn't disappear—it gets encoded into portable software that runs on any Kubernetes cluster, on any infrastructure provider, in any jurisdiction. Developer convenience, reimagined. Solid production databases, managed the GitOps way, on a European Kubernetes cloud. ## Follow the Money The total cloud spend will follow the same trajectory. Enterprises aren't going back to running their own data centers. The shift from capex to opex is permanent. The question isn't whether the money gets spent—it's where it flows. In the hyperscaler model, compute is a loss leader. The margins are in the managed services—the proprietary APIs that create lock-in and justify premium pricing. European cloud providers competing on compute alone were always fighting for scraps. The operator model redistributes this spend: **Commodities get priced as commodities.** Compute, storage, network—these become interchangeable. European providers like Scaleway, OVHcloud, and Hetzner compete just fine on infrastructure. When Kubernetes is the abstraction layer, the underlying provider matters less. Price, performance, jurisdiction—pick what matters to you. **The service layer goes to software.** The money that used to flow to AWS for RDS and SQS? It can flow to software development instead. To the companies and communities building operators. To the engineers encoding operational knowledge. If we play it right, this means open source—sustainable, foundation-governed projects that benefit everyone. **Investment stays local.** Instead of filling overseas coffers, we build our own digital muscle. European engineers, European companies, European open source communities—all strengthened by redirected cloud spend. ## The New Cloud Stack The future isn't one provider offering everything. It's a composable stack: **Layer 1: Kubernetes as the universal runtime.** Scaleway Kapsule, OVHcloud Managed Kubernetes, Exoscale SKS—they're all running the same API, the same ecosystem. The managed Kubernetes market is already commoditizing. **Layer 2: Operators as portable managed services.** Database? Deploy an operator. Message queue? Deploy an operator. Certificate management, secrets, identity? All operators. The software comes from specialists who focus on doing one thing exceptionally well. **Layer 3: Commodity external services.** DNS, email delivery, CDN, DDoS protection—these benefit from global scale but are also commodities with standard APIs. Use them without lock-in. **Layer 4: Specialist infrastructure.** GPU clusters for AI. High-memory instances for analytics. Edge nodes for low latency. Different providers excel at different things. Kubernetes lets you consume them all through a unified interface. This model inverts cloud economics. The generalist providers offer the compute foundation. Specialists offer differentiated capabilities. And operators—portable, open source, community-driven—deliver the managed service experience. No single provider controls the stack. No single jurisdiction controls the data. ## Why Europe Wins This Game The operator model plays to European strengths: **Open source is in our DNA.** European developers and companies are deeply embedded in the CNCF ecosystem. The operators powering this revolution—Flux, cert-manager, CloudNativePG, Strimzi—are community-built, foundation-governed. European values around collaborative development align naturally with this model. **We don't need hyperscaler scale.** A specialist provider with the best GPU infrastructure can compete on that strength alone. A regional provider with low-latency edge compute can serve local markets profitably. Kubernetes democratizes participation in the cloud economy. **Sovereignty becomes architecture.** When your "managed services" are operators running in your cluster, on your chosen provider, in your chosen jurisdiction—sovereignty isn't a feature request. It's the default. **Competition keeps everyone honest.** Operators are portable. If your provider doubles their prices, you migrate. The switching cost is infrastructure, not application rewriting. This pressure benefits customers and prevents the kind of lock-in that hyperscalers depend on. ## Accelerating the Transition The operator ecosystem is maturing fast, but gaps remain. Not every managed service has a production-ready operator equivalent. The developer experience isn't yet as smooth as a cloud console. Operational knowledge is still being encoded. This is where sharing accelerates everything. Every team that deploys operators in production and publishes their configurations reduces friction for the next team. Every contribution to operator projects improves the software for everyone. Every scaffold and tutorial that shows "here's how we run this on European infrastructure" builds the ecosystem. I've open-sourced a Kubernetes scaffold for Scaleway that embodies this approach: OpenTofu for infrastructure, Flux for GitOps, operators handling what they can. It's a starting point. Fork it. Improve it. Share what you build. ## The Revolution Is Already Here The next cloud revolution isn't about who builds the biggest data centers. It's about how software gets distributed and operated. When operational knowledge travels as code—as operators anyone can deploy—the hyperscaler managed service advantage evaporates. What remains is infrastructure: compute, storage, network. Commodities. And on that playing field, European providers compete just fine. The hyperscalers will adapt. They'll push proprietary Kubernetes extensions, try to lock in through AI services, find new vectors. But the fundamental shift toward portable, composable infrastructure favors diversity over monopoly. Europe doesn't need to build its own AWS. We tried that. It didn't work. Instead, we build the ecosystem that makes AWS optional. We redirect cloud spend from overseas lock-in to local capability. We turn the commodity trap into a commodity advantage. The Kubernetes operator revolution is that ecosystem. The rules have changed. Time to play a different game. --- *The Scaleway Project Scaffold is available at [github.com/aknostic/scaleway-project-scaffold](https://github.com/aknostic/scaleway-project-scaffold) under the Apache 2.0 license.* --- ## Building European Observability: Multi-Location Synthetic Monitoring with GitOps URL: https://clouds-of-europe.eu/content/practice/implementation-patterns/building-european-observability-multi-location-synthetic-monitoring-with-gitops Author: Jurg van Vliet Published: 2026-01-28 Category: Practice Type: Implementation Patterns When your platform serves a European audience, monitoring from US-East-1 tells you almost nothing useful. Network latency, regional routing policies, and even DNS resolution behave differently when your traffic doesn't cross the Atlantic. We learned this the hard way when our synthetic probes showed everything was fine while actual users in France reported occasional slow page loads. This article walks through how we built multi-location synthetic monitoring for Clouds of Europe—probing our endpoints every 30 seconds from Paris, Amsterdam, and Warsaw. But the interesting part isn't the monitoring itself. It's how we built it: infrastructure-as-code with Grafana Operator CRDs, consensus-based alerting that reduces false positives, and a testing procedure that uses Envoy Gateway SecurityPolicies to simulate real outages. ## Why European Vantage Points Matter Most managed monitoring services run their probes from a handful of locations, often concentrated in North America. This creates a blind spot: you're measuring performance through a lens that doesn't match your users' experience. For a European digital independence platform, this mismatch is both technical and philosophical. If your mission is European independence from US cloud infrastructure, it's incongruous to rely on American probe servers to tell you whether your site is up. We use Heystaq, a European managed observability platform built by Aknostic on Mimir, Loki and Tempo, which runs Blackbox Exporter instances in three European cities: | Location | Purpose | |----------|---------| | Paris (fr-par) | Western Europe, major internet exchange | | Amsterdam (nl-ams) | Northern Europe, AMS-IX proximity | | Warsaw (pl-waw) | Eastern Europe, emerging tech hub | This geographic distribution means we detect region-specific issues—a BGP misconfiguration affecting German traffic, a Scaleway datacenter hiccup in Paris—before they become widespread user complaints. ## The GitOps Pattern: Synthetic Targets as Configuration Traditional monitoring setups involve clicking through web UIs to add probe targets. This works until you need to reproduce your configuration after a disaster, audit who changed what, or promote changes through test environments. We define our synthetic targets in a ConfigMap that lives in Git: ```yaml apiVersion: v1 kind: ConfigMap metadata: name: synthetic-targets-clouds-of-europe labels: heystaq.com/synthetic-targets: "true" heystaq.com/tenant: "clouds-of-europe" data: targets.yaml: | - targets: ["placeholder"] labels: name: "coe-homepage" address: "https://clouds-of-europe.eu" module: "http_2xx" environment: "production" priority: "critical" __scrape_interval__: "30s" - targets: ["placeholder"] labels: name: "coe-health-api" address: "https://clouds-of-europe.eu/api/health" module: "http_2xx" environment: "production" priority: "critical" __scrape_interval__: "30s" - targets: ["placeholder"] labels: name: "coe-content-api" address: "https://clouds-of-europe.eu/api/content/cached?limit=1" module: "http_2xx" environment: "production" priority: "warning" __scrape_interval__: "30s" __probe_timeout__: "15s" ``` The `placeholder` in targets is a Heystaq convention—the actual target URL comes from the `address` label. What matters here is that every configuration decision is visible: the 30-second scrape interval, the 15-second timeout for the content API, the priority labels that determine alerting severity. When we need to add a new endpoint to monitor, we edit this file, commit it, and push. Flux picks up the change and applies it. No clicking, no screenshots of "how to configure monitoring," no wondering if production matches staging. ## Debugging the Paris DNS Problem A few weeks after deploying this setup, we noticed intermittent probe failures specifically from Paris. The homepage and health API were fine, but the content API endpoint was failing roughly 10% of the time—always from the French probe. The metrics told the story: `probe_duration_seconds` for successful probes was occasionally spiking to 9-10 seconds. Our default 10-second timeout wasn't leaving any margin. Digging deeper, we found the culprit: DNS resolution. The content API endpoint has a more complex hostname path that was triggering additional DNS lookups. From Paris, these lookups occasionally took 5+ seconds—possibly due to resolver congestion or routing to a distant upstream server. The fix was simple once we understood it: ```yaml __probe_timeout__: "15s" ``` This is the kind of regional quirk that US-based monitoring would never catch. Your American probes see 200ms DNS resolution consistently. Your French users see 5-second spikes that make your site feel broken. ## Consensus-Based Alerting The most valuable insight from multi-location monitoring isn't "is my site up?" but "is my site up *for most users*?" A single failed probe might indicate a problem with the probe itself, a regional network issue, or a transient blip. Our alerting rules use consensus: we only fire critical alerts when the majority of locations report failure. ```yaml - uid: af9j72o3pay2ob title: "EndpointDown" condition: C for: 1m labels: severity: "critical" annotations: description: "{{ $labels.instance }} is unreachable from majority of probe locations - likely service outage." summary: "{{ $labels.instance }} down - less than 50% of locations can reach it" data: - refId: A model: expr: avg by (instance, target_name, priority) (probe_success) ``` The PromQL expression `avg by (instance) (probe_success)` computes the average success rate across all probe locations. When this drops below 50%, we know the problem isn't localized—something is genuinely wrong. We also have a "LocationDown" alert for single-location failures, but it's severity "warning" rather than "critical." This gives us visibility into regional issues without waking someone at 3am for a Paris-specific network blip. ## Grafana Operator CRDs: Alerts as Code The traditional way to configure Grafana alerts involves the web UI: clicking through forms, hoping you don't fat-finger a threshold, and having no record of what changed. For a GitOps platform, this is untenable. We converted our entire alerting configuration to Grafana Operator Custom Resource Definitions. Here's what a synthetic monitoring alert group looks like: ```yaml apiVersion: grafana.integreatly.org/v1beta1 kind: GrafanaAlertRuleGroup metadata: name: synthetic-monitoring namespace: org-clouds-of-europe spec: instanceSelector: matchLabels: org: clouds-of-europe folderRef: synthetic-monitoring interval: 1m rules: - uid: cf9j72llgh0cgf title: "LocationDown" condition: C for: 2m labels: severity: "warning" annotations: runbook_url: "https://github.com/..." ``` This transformation was significant: 17 alert rules converted to 5 GrafanaAlertRuleGroup resources, 7 dashboards to GrafanaDashboard CRDs, plus notification policies and contact points. Everything lives in Git, everything is versioned, everything is reproducible. The `runbook_url` annotation is key. Every alert links directly to documentation explaining what the alert means and how to respond. When the pager goes off at 2am, you're not guessing—you're following a procedure. ## Testing Alerts Without Breaking Production How do you test your alerting pipeline without actually taking down production? We developed a procedure using Envoy Gateway SecurityPolicies. The idea is simple: block the probe IP addresses temporarily, verify alerts fire, then remove the block. But the implementation requires care—you don't want Flux to reconcile your temporary change away. ```bash # Suspend Flux first to prevent auto-reconciliation flux suspend helmrelease clouds-of-europe -n flux-system # Block the 3 probe IPs cat < 0) ) * 100 ``` If this ratio is consistently below 20%, you're over-provisioned. **CPU usage vs request**: ```promql # CPU usage as percentage of request ( rate(container_cpu_usage_seconds_total{container!="",container!="POD"}[5m]) / (kube_pod_container_resource_requests{resource="cpu"} > 0) ) * 100 ``` CPU is trickier because usage is spiky. Look at 95th percentile over a week, not instant values. ## Setting Requests Based on Reality **Step 1**: Query actual usage over at least a week (preferably 30 days) **Step 2**: Find p95 or p99 usage (not average—you need to handle peaks) **Step 3**: Add headroom: - Memory: +30-50% (memory usage spikes matter) - CPU: +50-100% (CPU is burstable, be generous) **Step 4**: Set limits: - Memory limit: 2x request (catch runaway processes, allow temporary spikes) - CPU limit: 2-4x request (or no limit—CPU throttling hurts performance) **Example**: ```yaml # Before (guessed) resources: requests: memory: "2Gi" cpu: "1000m" limits: memory: "4Gi" cpu: "2000m" # After (measured) # p95 usage: 180MB memory, 120m CPU resources: requests: memory: "256Mi" # 180MB * 1.4 ≈ 256MB cpu: "200m" # 120m * 1.6 ≈ 200m limits: memory: "512Mi" # 2x request cpu: "1000m" # 5x request (allow burst) ``` This is 8x better memory density. Instead of 16 pods per 32GB node, you can fit 120 pods. ## Quarterly Review Process Resource usage changes over time. Features are added, traffic patterns shift. Review quarterly: ```bash # Query Prometheus for usage over last 30 days # Generate recommendations # Apply changes gradually ``` Make changes incrementally. Update one service, monitor for a week, then move to the next. If you see OOMKills, you were too aggressive—add more headroom. ## Spot Instances for Appropriate Workloads Not all workloads need guaranteed capacity. Some can tolerate interruption: **CI/CD pipelines**: Can restart if preempted. Use spot instances, save 60-80%. **Batch processing**: Can checkpoint and resume. Perfect for spot. **Development environments**: Interruption is annoying, not critical. We run dev on spot. **Not appropriate for spot**: - User-facing production applications (interruption affects users) - Stateful databases (complex to handle interruption safely) - Real-time processing (can't tolerate delays) ## Data Hygiene for Storage Efficiency Storage seems cheap, so data accumulates. Logs, backups, old snapshots—it all adds up. **Define retention policies early**: ```yaml # Object storage lifecycle lifecycle_rule: enabled: true expiration: days: 365 transition: days: 90 storage_class: GLACIER ``` **Actually delete when retention expires**. We saved 40% storage costs by implementing retention policies and sticking to them. ## Why Efficiency Matters **Cost**: Right-sizing our infrastructure saved roughly €300/month. That's €3,600/year. **Performance**: After right-sizing, we had more available capacity for bursts. Counter-intuitively, using resources more efficiently improved performance. **Environment**: Fewer idle resources mean less energy wasted. Data centers consume about 1.5% of global electricity. Efficiency at scale matters. Efficiency isn't just environmental virtue. Efficient systems cost less, run better, and are easier to operate. **Sources:** - [Kubernetes Resource Management](https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/) - [Prometheus Queries for Right-Sizing](https://www.robustperception.io/existential-issues-with-metrics) #efficiency #rightsizing #kubernetes #sustainability #costoptimization --- ## Your First Migration: From Zero to Production in Weeks URL: https://clouds-of-europe.eu/content/practice/getting-started-guides/your-first-migration-from-zero-to-production-in-weeks Author: Jurg van Vliet Published: 2025-11-28 Category: Practice Type: Getting Started Guides Tags: gitops, infrastructure, kubernetes, migration, stepbystep ## Week-by-Week Timeline Moving your first workload to European infrastructure in four weeks. This is realistic for a stateless application with an experienced team. Adjust the timeline based on complexity and organisational constraints. ## Week 1: Infrastructure Setup **Day 1-2: Provider selection and access** Choose your provider based on requirements (see our evaluation criteria article). Create account, set up billing, configure initial access. For this guide, we'll use Scaleway, but the patterns apply to OVHcloud, Hetzner, or IONOS. ```bash # Install tools brew install opentofu kubectl flux age sops # Configure Scaleway CLI scw init ``` **Day 3-5: Infrastructure as Code** Create OpenTofu configuration for your cluster: ```hcl # main.tf terraform { required_providers { scaleway = { source = "scaleway/scaleway" version = "~> 2.0" } } } variable "cluster_name" { default = "pilot-cluster" } resource "scaleway_k8s_cluster" "main" { name = var.cluster_name version = "1.28.5" cni = "cilium" region = "fr-par" autoscaler_config { disable_scale_down = false scale_down_delay_after_add = "5m" } } resource "scaleway_k8s_pool" "main" { cluster_id = scaleway_k8s_cluster.main.id name = "main-pool" node_type = "DEV1-M" size = 3 autoscaling = true min_size = 3 max_size = 6 } output "kubeconfig" { value = scaleway_k8s_cluster.main.kubeconfig[0].config_file sensitive = true } ``` Apply the configuration: ```bash tofu init tofu plan tofu apply # Get kubeconfig tofu output -raw kubeconfig > kubeconfig.yaml export KUBECONFIG=$(pwd)/kubeconfig.yaml # Verify cluster access kubectl get nodes ``` ## Week 2: Application Deployment **Day 1-2: Prepare application manifests** If your application isn't yet in Kubernetes format, containerise and create manifests: ```yaml # deployment.yaml apiVersion: apps/v1 kind: Deployment metadata: name: myapp namespace: default spec: replicas: 2 selector: matchLabels: app: myapp template: metadata: labels: app: myapp spec: containers: - name: app image: your-registry.com/myapp:v1.0.0 ports: - containerPort: 8080 resources: requests: memory: "256Mi" cpu: "100m" limits: memory: "512Mi" cpu: "500m" env: - name: DATABASE_URL valueFrom: secretKeyRef: name: app-secrets key: database-url --- apiVersion: v1 kind: Service metadata: name: myapp spec: selector: app: myapp ports: - port: 80 targetPort: 8080 ``` **Day 3-4: Test deployment manually** Deploy and verify functionality: ```bash # Create namespace kubectl create namespace myapp # Deploy application kubectl apply -f deployment.yaml -n myapp # Check status kubectl get pods -n myapp kubectl logs -f deployment/myapp -n myapp # Test locally kubectl port-forward svc/myapp 8080:80 -n myapp curl http://localhost:8080 ``` **Day 5: Handle secrets** Generate age key for encryption: ```bash # Generate age keypair age-keygen -o age.key # Public key for .sops.yaml: age1ql3z7hjy54pw3hyww5ayyfg7zqgvc7w3j2elw8zmrj2kg5sfn9aqmcac8p # Private key stored in: age.key ``` Create and encrypt secrets: ```bash # Create plaintext secret cat > secrets.yaml < .sops.yaml < secrets.enc.yaml # Commit encrypted version only git add secrets.enc.yaml .sops.yaml git commit -m "Add encrypted application secrets" # Delete plaintext (important!) rm secrets.yaml ``` ## Week 3: GitOps Setup **Day 1-2: Bootstrap Flux** ```bash # Install Flux controllers flux bootstrap github \ --owner=your-org \ --repository=your-repo \ --path=gitops/clusters/pilot \ --personal=false # Create SOPS decryption secret kubectl create secret generic sops-age \ --namespace=flux-system \ --from-file=age.agekey=age.key ``` **Day 3-4: Migrate manifests to GitOps** Create repository structure: ``` gitops/ ├── clusters/ │ └── pilot/ │ └── flux-system/ ├── infrastructure/ │ └── pilot/ │ └── cert-manager/ └── apps/ └── pilot/ ├── myapp/ │ ├── deployment.yaml │ ├── service.yaml │ └── secrets.enc.yaml └── kustomization.yaml ``` Create Flux Kustomization: ```yaml # gitops/clusters/pilot/apps.yaml apiVersion: kustomize.toolkit.fluxcd.io/v1 kind: Kustomization metadata: name: apps namespace: flux-system spec: interval: 10m path: ./gitops/apps/pilot prune: true sourceRef: kind: GitRepository name: flux-system decryption: provider: sops secretRef: name: sops-age ``` Commit and push: ```bash git add gitops/ git commit -m "Add GitOps structure and application manifests" git push # Watch Flux reconcile flux get kustomizations --watch ``` **Day 5: Verify GitOps workflow** Test the full cycle: ```bash # Make a change in git # Edit gitops/apps/pilot/myapp/deployment.yaml # Change replica count from 2 to 3 git commit -am "Scale myapp to 3 replicas" git push # Watch Flux apply the change flux reconcile kustomization apps --with-source kubectl get pods -n myapp -w ``` ## Week 4: Testing and Cutover **Day 1-2: Comprehensive testing** Run your full test suite against the new deployment: ```bash # Port forward for testing kubectl port-forward svc/myapp 8080:80 -n myapp # Run integration tests API_BASE_URL=http://localhost:8080 npm test # Load testing (if applicable) # Performance comparison with old environment ``` **Day 3: Traffic split preparation** Set up Gateway API for gradual traffic shifting (if moving from existing deployment): ```yaml apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: name: myapp-route spec: parentRefs: - name: main-gateway rules: - matches: - path: type: PathPrefix value: / backendRefs: - name: myapp-old # Existing deployment port: 80 weight: 90 - name: myapp # New deployment port: 80 weight: 10 # Start with 10% traffic ``` **Day 4: Execute cutover** Gradually shift traffic: ```bash # Hour 0: 10% new, 90% old (already set) # Monitor error rates, latency, logs # Hour 2: 50% new, 50% old (if no issues) # Update HTTPRoute weights # Hour 4: 100% new, 0% old (if still no issues) # Update HTTPRoute to remove old backend # Monitor for 24 hours before removing old infrastructure ``` **Day 5: Documentation and cleanup** Document what you learned: - What worked well - What was harder than expected - What you'd do differently - Actual vs estimated timeline Clean up old infrastructure once stable. ## Common Pitfalls **Underestimating DNS propagation**: Allow time for DNS changes to propagate (up to 48 hours for some zones, though usually much faster). **Missing monitoring**: Set up monitoring **before** cutover, not after. You need visibility into the new environment. **Insufficient testing**: Test failure scenarios, not just happy paths. What happens when the database is unreachable? **Secrets in logs**: Review logs to ensure secrets aren't being logged. This is an easy mistake to make. **No rollback plan**: Know exactly how to switch back to the old environment if needed. ## Success Criteria You've succeeded when: - Application runs stable on new infrastructure - All tests pass - Traffic is served with acceptable latency - Monitoring shows no elevated error rates - GitOps workflow is working (changes via git) - Team knows how to debug and make changes - Documentation is complete Then celebrate, review lessons learned, and plan the next migration. #migration #kubernetes #gitops #stepbystep #infrastructure --- ## AI as Learning Accelerator: Modern Tooling for Small Teams URL: https://clouds-of-europe.eu/content/practice/tools-templates/ai-as-learning-accelerator-modern-tooling-for-small-teams Author: Jurg van Vliet Published: 2025-11-25 Category: Practice Type: Tools & Templates Tags: ai, development, learning, productivity, tooling ## What AI Actually Did for This Project This platform—Clouds of Europe—was built in six months by a small team (mostly solo, with occasional help). That timeline included: - Multi-cluster Kubernetes setup (local, test, production) - GitOps with Flux v2 and SOPS encryption - OpenTofu for infrastructure provisioning - Complete observability stack (Prometheus, Loki, Grafana) - Next.js application with authentication, content system, and event management - 264 API tests with 100% endpoint coverage - Playwright E2E tests - Production deployment on Scaleway Six months isn't impressive by Silicon Valley standards. For infrastructure projects, it's unusually fast. AI made the difference. ## What AI Accelerated **1. Learning complex systems faster** Kubernetes Gateway API was new to all of us. Instead of reading documentation for hours, we could ask: "How do I configure an HTTPRoute with path-based routing and header manipulation?" Claude provided working examples with explanations. We learned by doing, with immediate feedback. **2. Boilerplate and scaffolding** OpenTofu modules have a lot of boilerplate—variables, outputs, provider configuration. AI can generate the structure in seconds: "Create an OpenTofu module for a Scaleway Kapsule cluster with configurable node pools and autoscaling." We review, adjust, test. But the scaffolding is instant. **3. Debugging** When Flux reconciliation failed with cryptic errors, we could paste the full error and context: "Flux shows 'kustomize build failed' with this error: [paste]. Here's my kustomization.yaml: [paste]. What's wrong?" AI spots issues humans miss—indentation errors, missing fields, incompatible API versions. **4. Documentation and explanation** Complex Kubernetes manifests with CRDs, policies, and networking are hard to read. AI can explain what a manifest actually does: "Explain this Gateway API configuration: [paste]" This is learning tool. You don't just copy-paste and hope—you understand what you're deploying. ## What AI Didn't Replace **Architecture decisions**: We chose Kubernetes, Flux, Scaleway, Gateway API, Next.js. AI provided information about options, but these were human decisions based on project requirements. **Domain expertise**: AI doesn't know your business, your compliance requirements, your performance constraints. We decided what needed to be built. **Critical judgment**: When AI suggests three approaches, you need to evaluate tradeoffs. AI can explain pros and cons; it can't decide what matters for your context. **Debugging complex issues**: AI helps with first-level debugging—syntax errors, missing configuration, common mistakes. Deep system issues still require understanding the stack. **Code review**: AI-generated code needs review. Sometimes it's subtly wrong—syntactically correct but semantically flawed. Human review catches this. ## Effective AI Use Patterns **Start with understanding**: Before asking AI to generate code, understand what you're trying to achieve. AI amplifies your intent; unclear intent produces unclear results. **Iterate and refine**: First response is rarely perfect. Refine the prompt, add context, specify constraints. AI conversation is iterative. **Verify everything**: Test AI-generated code. Don't assume it works. Our test coverage exists partly to catch AI mistakes. **Learn from the output**: Don't just copy-paste. Read the generated code, understand why it's structured that way, learn the patterns. ## Honest Assessment: 6 Months vs What? Without AI, this project would have taken: - **12-18 months with experienced team**: Someone who already knows Kubernetes, Flux, Gateway API, and the ecosystem could build this in a year to a year and a half. - **24+ months learning from scratch**: Learning all these technologies from documentation, experimenting, debugging—that's a multi-year journey. AI compressed learning time. We learned Kubernetes Gateway API, Flux GitOps patterns, OpenTofu module structure, and Playwright testing in weeks instead of months. The concepts still had to be learned—AI just made the learning path shorter. ## What This Means for Small Teams **You can build more with less**: Projects that required dedicated DevOps engineers are now accessible to small teams. The cognitive load of learning new technology has decreased. **You need different skills**: Less "memorise syntax," more "design good systems and critically evaluate solutions." The judgment layer becomes more important. **The pace of change accelerates**: When learning new tools is faster, you can adopt new technologies more readily. This can be good (better tools) or bad (tool churn). ## The Tool, Not the Author We document AI's role honestly. This doesn't diminish the work—the architecture, judgment, testing, refinement, and actual building were human. AI was a powerful assistant that compressed learning cycles. Think of AI as a very knowledgeable pair programmer who: - Never gets tired - Knows a broad range of technologies - Doesn't judge stupid questions - Can't make final decisions - Sometimes confidently suggests wrong approaches Used well, it's transformative. Used poorly, it produces plausible-looking broken code. #ai #tooling #learning #development #productivity --- ## Gateway API in Practice: The Nginx→Envoy Migration URL: https://clouds-of-europe.eu/content/policy-sovereignty/open-source-standards/gateway-api-in-practice-the-nginxenvoy-migration Author: Jurg van Vliet Published: 2025-11-21 Category: Policy & Sovereignty Type: Open Source & Standards Tags: envoy, gatewayapi, kubernetes, migration, standards ## The Retirement Announcement November 2025: Kubernetes SIG Network and the Security Response Committee announced the retirement of Ingress NGINX. Best-effort maintenance would continue until March 2026. After that: no releases, no bugfixes, no security updates. The reasons were clear: keeping ingress-nginx aligned across Kubernetes versions, NGINX releases, and Helm charts had become unsustainable. Recent high-severity vulnerabilities (including one allowing complete cluster takeover) exposed how heavy the maintenance burden had become. We'd been running Nginx Ingress for years. It worked fine. But retirement meant migration was now urgent—continuing on unmaintained software wasn't acceptable for production. ## Why Gateway API Made Migration Straightforward Gateway API is Kubernetes' successor to Ingress. The key architectural improvement: **separation of intent from implementation**. **With Ingress (the old way):** Routing configuration was tightly coupled to the ingress controller implementation. Switching from Nginx to Traefik or HAProxy meant rewriting configurations because each controller used different annotations. Example: ```yaml # Nginx Ingress apiVersion: networking.k8s.io/v1 kind: Ingress metadata: annotations: nginx.ingress.kubernetes.io/rewrite-target: / nginx.ingress.kubernetes.io/ssl-redirect: "true" spec: ingressClassName: nginx rules: - host: api.example.com http: paths: - path: / backend: service: name: api port: 8080 ``` These annotations are Nginx-specific. Moving to different ingress controller requires rewriting. **With Gateway API (the new way):** Routing intent is expressed in standard resources (HTTPRoute). Implementation details live in separate Gateway resources. Changing implementations doesn't require touching routing configuration. Example: ```yaml # HTTPRoute (controller-agnostic) apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: name: api-route spec: parentRefs: - name: main-gateway hostnames: - "api.example.com" rules: - matches: - path: type: PathPrefix value: / backendRefs: - name: api port: 8080 --- # Gateway (implementation-specific) apiVersion: gateway.networking.k8s.io/v1 kind: Gateway metadata: name: main-gateway spec: gatewayClassName: envoy # Could be cilium, istio, etc. listeners: - name: https protocol: HTTPS port: 443 tls: mode: Terminate certificateRefs: - name: api-tls ``` The HTTPRoute is implementation-agnostic. It works with Envoy, Cilium, Istio, or any Gateway API-compliant controller. Only the Gateway resource knows which implementation you're using. ## Our Migration Process **Week 1: Deploy Envoy Gateway alongside Nginx** Both controllers running. Nginx handling all traffic. Envoy Gateway deployed but not routing anything yet. ```bash # Install Envoy Gateway helm install envoy-gateway oci://docker.io/envoyproxy/gateway-helm \ --version v1.2.8 \ --namespace envoy-gateway-system \ --create-namespace # Verify both controllers running kubectl get pods -n ingress-nginx kubectl get pods -n envoy-gateway-system ``` **Week 2: Create Gateway and HTTPRoutes** Convert Nginx Ingress resources to Gateway API: ```yaml # Gateway (infrastructure layer) apiVersion: gateway.networking.k8s.io/v1 kind: Gateway metadata: name: production-gateway namespace: infrastructure spec: gatewayClassName: envoy listeners: - name: https protocol: HTTPS port: 443 hostname: "*.clouds-of-europe.eu" tls: mode: Terminate certificateRefs: - name: wildcard-tls namespace: cert-manager --- # ReferenceGrant (allows Gateway to reference cert in different namespace) apiVersion: gateway.networking.k8s.io/v1beta1 kind: ReferenceGrant metadata: name: allow-gateway-cert-access namespace: cert-manager spec: from: - group: gateway.networking.k8s.io kind: Gateway namespace: infrastructure to: - group: "" kind: Secret ``` ```yaml # HTTPRoute (application layer) apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: name: app-route namespace: app spec: parentRefs: - name: production-gateway namespace: infrastructure hostnames: - "clouds-of-europe.eu" - "www.clouds-of-europe.eu" rules: - backendRefs: - name: app port: 3000 ``` **Week 3: Gradual traffic shift** Update DNS to point to Envoy Gateway load balancer. Monitor closely: - Error rates (should be unchanged) - Latency (should be comparable) - TLS handshake success - Certificate validation If issues appear, DNS can revert to Nginx immediately. **Week 4: Monitor new system** Envoy handling 100% traffic. Nginx still deployed but receiving nothing. Watch for: - Any edge cases not handled by Gateway API - Performance characteristics - Resource usage of new controller **Week 5: Remove Nginx** After week of stable Envoy operation, remove Nginx Ingress completely: ```bash helm uninstall ingress-nginx -n ingress-nginx kubectl delete namespace ingress-nginx ``` Total migration: 5 weeks. Zero incidents. Zero downtime. That's the goal. ## The Boring Change Principle If changing a core component requires heroism, your architecture has a problem. Good architecture makes change boring. Routine. Unremarkable. We don't celebrate impressive migrations—we celebrate migrations so smooth nobody notices. Gateway API enabled this. By separating routing intent (HTTPRoute) from implementation (Gateway), we could swap the proxy layer without touching application configuration. ## What We Learned **Gateway API works:** The abstraction is sound. HTTPRoute is expressive enough for complex routing while remaining implementation-agnostic. **ReferenceGrants are essential:** Without ReferenceGrants, you're forced to copy TLS certificates between namespaces. This creates synchronization problems and duplicates sensitive data. ReferenceGrants solve this cleanly. **Not all features are standardized yet:** We needed some Envoy-specific features (BackendTrafficPolicy for circuit breaking). These required Envoy Gateway CRDs—not portable. But 90% of our routing is standard HTTPRoute. **Separation of concerns works:** Platform team owns Gateways (infrastructure). Application teams own HTTPRoutes (routing). Neither needs elevated permissions for the other's layer. **Implementation matters:** Gateway API is a standard. Implementations (Envoy, Cilium, Istio) have different performance characteristics, feature sets, and maturity levels. The standard enables choice; you still need to choose wisely. ## Looking Forward Gateway API is now GA (generally available) in Kubernetes. Ingress is effectively deprecated with Nginx's retirement. This is the pattern for infrastructure evolution: **new standard emerges, provides better abstractions, gradual migration from old to new**. Organizations building on Gateway API today are building on the future of Kubernetes networking. When the next generation of proxy technology emerges—and it will—Gateway API will enable migration without touching application routing. **The pattern:** Own the interface, not the implementation. Standards provide interfaces. Products provide implementations. Build on standards. **Sources:** - [Ingress NGINX Retirement Announcement (Nov 2025)](https://kubernetes.io/blog/2025/11/11/ingress-nginx-retirement/) - [InfoQ: Kubernetes Community Retires Ingress NGINX](https://www.infoq.com/news/2025/11/kubernetes-ingress-nginx/) - [Gateway API Documentation](https://gateway-api.sigs.k8s.io/) #gatewayapi #envoy #kubernetes #migration #standards --- ## Multi-Cluster Networking: mTLS with Gateway API URL: https://clouds-of-europe.eu/content/practice/expert-insights/multicluster-networking-mtls-with-gateway-api Author: Jurg van Vliet Published: 2025-11-17 Category: Practice Type: Expert Insights Tags: gatewayapi, kubernetes, mtls, networking, security ## The Cross-Cluster Requirement We run multiple Kubernetes clusters: local (Kind), test (Scaleway), and production (Scaleway). Each cluster has Prometheus scraping local metrics. For centralised monitoring, our management cluster runs Grafana and Mimir. Grafana needs to query Prometheus in each cluster. This is cross-cluster networking, and it needs to be secure. Requirements: - Grafana queries Prometheus in test/production clusters - Connections must be encrypted (TLS) - Connections must be mutually authenticated (client certs) - No public exposure of Prometheus endpoints - Manageable certificate lifecycle Gateway API with cert-manager solves this cleanly. ## The Gateway API Pattern We use **separate Gateways** for public and internal traffic: **Public Gateway**: Serves application traffic, uses Let's Encrypt certificates, no client authentication required. **Internal Gateway**: Serves cluster-to-cluster traffic, requires mutual TLS (both server and client certs), not publicly accessible. This separation means internal endpoints aren't accidentally exposed. Different Gateway, different TLS configuration, different security posture. ## Implementation with cert-manager **Step 1: Create cluster CA** Each cluster has a cert-manager ClusterIssuer for internal certificates: ```yaml apiVersion: cert-manager.io/v1 kind: ClusterIssuer metadata: name: internal-ca spec: ca: secretName: internal-ca-secret ``` The CA certificate is generated once and stored in `internal-ca-secret`. This CA issues certificates for both servers (Prometheus endpoints) and clients (Grafana). **Step 2: Server certificate for Prometheus** ```yaml apiVersion: cert-manager.io/v1 kind: Certificate metadata: name: prometheus-server-cert namespace: monitoring spec: secretName: prometheus-tls issuerRef: name: internal-ca kind: ClusterIssuer dnsNames: - prometheus.monitoring.svc.cluster.local - prometheus.example.com ``` cert-manager creates `prometheus-tls` secret with certificate and key. **Step 3: Client certificate for Grafana** ```yaml apiVersion: cert-manager.io/v1 kind: Certificate metadata: name: grafana-client-cert namespace: monitoring spec: secretName: grafana-client-tls issuerRef: name: internal-ca kind: ClusterIssuer commonName: grafana-client usages: - client auth ``` **Step 4: Gateway configuration for mTLS** This is where Gateway API shows its value: ```yaml apiVersion: gateway.networking.k8s.io/v1 kind: Gateway metadata: name: internal-gateway namespace: monitoring spec: gatewayClassName: envoy-gateway listeners: - name: prometheus-mtls protocol: HTTPS port: 9090 hostname: "prometheus.example.com" tls: mode: Terminate certificateRefs: - name: prometheus-tls kind: Secret ``` For client validation, we use Envoy Gateway's BackendTLSPolicy (Gateway API extension): ```yaml apiVersion: gateway.envoyproxy.io/v1alpha1 kind: ClientTrafficPolicy metadata: name: mtls-validation namespace: monitoring spec: targetRef: group: gateway.networking.k8s.io kind: Gateway name: internal-gateway tls: clientValidation: caCertificateRefs: - name: internal-ca-secret group: "" kind: Secret ``` This configuration: - Terminates TLS with the server certificate - Requires client certificate for authentication - Validates client cert against the cluster CA **Step 5: HTTPRoute to Prometheus** ```yaml apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: name: prometheus-route namespace: monitoring spec: parentRefs: - name: internal-gateway namespace: monitoring hostnames: - "prometheus.example.com" rules: - matches: - path: type: PathPrefix value: / backendRefs: - name: prometheus port: 9090 ``` ## ReferenceGrants for Cross-Namespace Access Gateway API enforces security: a Gateway in namespace A can't reference a Secret in namespace B without explicit permission. **Problem**: Gateway is in `monitoring` namespace. CA certificate might be in `cert-manager` namespace. **Solution**: ReferenceGrant ```yaml apiVersion: gateway.networking.k8s.io/v1beta1 kind: ReferenceGrant metadata: name: allow-gateway-to-cert-manager namespace: cert-manager spec: from: - group: gateway.networking.k8s.io kind: Gateway namespace: monitoring to: - group: "" kind: Secret ``` This grants the Gateway in `monitoring` namespace permission to reference Secrets in `cert-manager` namespace. Explicit, auditable, secure. ## What We Learned **Separation is cleaner than complex policies**: We initially tried to use one Gateway with different TLS policies per route. This was complicated and error-prone. Separate Gateways—one public, one internal—is simpler. **ReferenceGrants add friction (intentionally)**: Having to explicitly grant cross-namespace access feels like extra configuration. But it prevents accidental exposure. The friction is security. **Certificate rotation is automatic**: cert-manager handles renewal. Certificates rotate before expiry without manual intervention. This is better than static certificates that require runbooks. **Debugging TLS is still hard**: When mTLS doesn't work, error messages are often cryptic. We added detailed logging and tested thoroughly in test environment before production. **Gateway API abstracts the proxy**: We run Envoy Gateway, but the HTTPRoute definitions don't depend on Envoy specifics (except TLS validation configuration). If we need to switch to Istio or Cilium later, most configuration remains unchanged. **Sources:** - [Gateway API documentation](https://gateway-api.sigs.k8s.io/) - [cert-manager documentation](https://cert-manager.io/) - [Envoy Gateway mTLS guide](https://gateway.envoyproxy.io/docs/) #gatewayapi #mtls #networking #kubernetes #security --- ## GitOps as Source of Truth: Rebuilding Clusters from Git URL: https://clouds-of-europe.eu/content/practice/implementation-patterns/gitops-as-source-of-truth-rebuilding-clusters-from-git Author: Jurg van Vliet Published: 2025-11-15 Category: Practice Type: Implementation Patterns Tags: disasterrecovery, flux, gitops, kubernetes, sourceoftruth ## What Source of Truth Actually Means Kubernetes workloads are disposable—when a node fails, the orchestrator reschedules elsewhere. But this only works if the *state* is stored somewhere reliable. What was running on that node? What configuration did it have? With GitOps, the answer is always: whatever's in the git repository. Source of truth means: - The repository defines what **should** be running - The cluster continuously reconciles to match - Divergence is automatically corrected - You can rebuild from the repository alone We've tested this. Not in a theoretical drill—in actual cluster rebuilds. It works. ## The Rebuild Process (Tested) We've done this multiple times when setting up new environments and testing disaster recovery: **Step 1: Provision infrastructure** (~15 minutes) ```bash cd infrastructure/opentofu/environments/production tofu init tofu apply -auto-approve # This creates: # - Kubernetes cluster (Scaleway Kapsule) # - Node pools (3 nodes, multi-AZ) # - Object storage buckets # - DNS records # - Load balancers ``` **Step 2: Bootstrap Flux** (~5 minutes) ```bash # Get cluster credentials export KUBECONFIG=./kubeconfig-production.yaml # Bootstrap Flux flux bootstrap github \ --owner=aknostic \ --repository=clouds-of-europe \ --path=gitops/clusters/production \ --personal # Create SOPS decryption key kubectl create secret generic sops-age \ --namespace=flux-system \ --from-file=age.agekey=$HOME/.config/sops/age/production.key ``` Flux installs itself and connects to the Git repository. **Step 3: Wait for reconciliation** (~30-40 minutes) ```bash # Watch Flux sync everything flux get kustomizations --watch # Monitor pods coming up watch kubectl get pods --all-namespaces ``` Flux automatically deploys (in order): 1. Infrastructure components (cert-manager, external-dns) 2. Monitoring stack (Prometheus, Grafana, Loki) 3. Application namespaces and RBAC 4. Application deployments 5. Gateway API routes and certificates 6. Secrets (decrypted from SOPS) **Step 4: Verify** (~5 minutes) ```bash # Check all pods running kubectl get pods --all-namespaces # Check ingress working curl https://clouds-of-europe.eu # Check database kubectl exec -it postgres-cluster-1 -n app -- \ psql -U postgres -c "SELECT COUNT(*) FROM users;" ``` **Total time: ~60 minutes** from "cluster doesn't exist" to "serving production traffic." No runbooks to follow manually. No configurations to remember. No tribal knowledge required. Just: provision infrastructure, bootstrap Flux, wait. ## Why This Matters **Confidence**: Knowing you CAN rebuild eliminates a category of anxiety. Infrastructure is cattle, not pets. Lose a cluster? Rebuild it. **Disaster recovery**: Real DR requires testing. We've tested this. Multiple times. In different environments. It works. **Documentation**: Git IS the documentation. Want to know how Grafana is configured? Read the Helm values in Git. Want to know what version of PostgreSQL? Read the manifest. **Onboarding**: New team member: "Clone the repo, read the gitops/ directory." That's the entire system, readable and navigable. ## Preventing Drift Drift is when cluster state diverges from Git state. This happens through: - Manual `kubectl apply` commands - UI changes (clicking buttons in dashboards) - Scripts that modify resources directly - "Quick fixes" during incidents **Why drift is dangerous:** 1. **Git lies about reality**: Documentation says X, cluster runs Y. Which is truth? 2. **Rebuilds fail**: Rebuild from Git produces different result than current cluster 3. **Changes get lost**: Someone fixes something manually, then cluster update reverts it 4. **Debugging is impossible**: Logs show configuration that doesn't match Git **Preventing drift: Flux with prune: true** ```yaml apiVersion: kustomize.toolkit.fluxcd.io/v1 kind: Kustomization metadata: name: apps namespace: flux-system spec: prune: true # Delete resources not in Git path: ./gitops/apps/production sourceRef: kind: GitRepository name: flux-system ``` Resources not defined in Git get deleted at next reconciliation. This sounds aggressive. It enforces discipline. **Result:** - Manual `kubectl apply`? Resource gets deleted at next reconciliation - UI changes? Reverted within minutes - Everything stays consistent with Git ## The No kubectl apply Rule We have a simple rule: **No `kubectl apply` against production. Ever.** Changes go through Git: ```bash # Wrong kubectl apply -f hotfix.yaml # Don't do this # Right git add hotfix.yaml git commit -m "Fix: increase memory limit for api-server" git push # Wait for Flux to reconcile ``` For urgent changes: ```bash # Trigger immediate reconciliation flux reconcile kustomization apps --with-source ``` This makes "change via Git" acceptable even during incidents. Commit fix, trigger reconciliation, change is live in ~30 seconds. ## Our Repository Structure ``` clouds-of-europe/ ├── infrastructure/ │ └── opentofu/ │ └── environments/ │ ├── foundation/ # S3, registry, DNS │ ├── management/ # Monitoring cluster │ ├── test/ # Test environment │ └── production/ # Production └── gitops/ └── clusters/ ├── management/ │ ├── flux-system/ # Flux components │ ├── infrastructure/ # cert-manager, external-dns │ └── observability/ # Prometheus, Grafana, Loki ├── test/ │ ├── flux-system/ │ ├── infrastructure/ │ └── apps/ # Application deployments └── production/ ├── flux-system/ ├── infrastructure/ └── apps/ ``` Every resource defined in Git. Nothing manual. Nothing in someone's head. ## Practical Benefits **Reproducibility**: Same Git commit = same infrastructure state. Testing, staging, production can be identical. **Auditability**: "Who changed the firewall rules?" `git log` shows exactly who, when, why. **Review process**: Infrastructure changes go through pull requests. Second pair of eyes before production. **Rollback**: Bad deployment? `git revert` and wait for reconciliation. Clean, auditable rollback. **Sources:** - [Flux documentation](https://fluxcd.io/) - [GitOps Principles](https://opengitops.dev/) #gitops #flux #sourceoftruth #disasterrecovery #kubernetes --- ## SOPS and AGE: Practical Secrets Management in GitOps URL: https://clouds-of-europe.eu/content/practice/implementation-patterns/sops-and-age-practical-secrets-management-in-gitops Author: Jurg van Vliet Published: 2025-11-12 Category: Practice Type: Implementation Patterns Tags: age, gitops, secrets, security, sops ## The Secrets in Git Problem GitOps means infrastructure as code, versioned in git. But what about secrets—database passwords, API keys, TLS certificates? The obvious answer: don't commit secrets to git. The practical problem: how do you deploy them? Manual secret management doesn't scale, breaks reproducibility, and creates drift. SOPS (Secrets OPerationS) solves this. You can commit secrets to git—encrypted—while maintaining the GitOps workflow. ## How SOPS Works SOPS encrypts the values in YAML or JSON files while leaving the structure readable: ```yaml # secrets.enc.yaml (encrypted with SOPS) apiVersion: v1 kind: Secret metadata: name: database-credentials namespace: production stringData: username: postgres password: ENC[AES256_GCM,data:Qr8yLKJ...,iv:abc123...,tag:xyz789...,type:str] database: clouds_of_europe sops: kms: [] age: - recipient: age1ql3z7hjy54pw3hyww5ayyfg7zqgvc7w3j2elw8zmrj2kg5sfn9aqmcac8p enc: | -----BEGIN AGE ENCRYPTED FILE----- YWdlLWVuY3J5cHRpb24ub3JnL3YxCi0+IFgyNTUxOSB... -----END AGE ENCRYPTED FILE----- ``` The structure is visible. You can see it's a Secret, see the keys (username, password, database), but the values are encrypted. Only systems with the decryption key can read the actual values. ## AGE Keys for Encryption We use age (Actually Good Encryption) for key management. Age is simpler than GPG—no key servers, no web-of-trust complexity, just a keypair. **Generating keys**: ```bash # Generate age keypair age-keygen -o age.key # Public key (for encryption): age1ql3z7hjy54pw3hyww5ayyfg7zqgvc7w3j2elw8zmrj2kg5sfn9aqmcac8p # Private key (for decryption): stored in age.key ``` Store the private key somewhere secure. For local development, that might be `~/.config/sops/age/keys.txt`. For CI/CD, GitHub secrets or equivalent. For Flux in-cluster, a Kubernetes secret. ## Configuration with .sops.yaml At the repository root, `.sops.yaml` defines encryption rules: ```yaml creation_rules: # Test environment secrets - path_regex: gitops/infrastructure/test/.*.yaml$ age: age1test...public-key-here # Production environment secrets - path_regex: gitops/infrastructure/production/.*.yaml$ age: age1prod...public-key-here ``` Different environments use different keys. If the test key is compromised, production secrets remain safe. This is per-environment isolation—the same principle as separate kubeconfig files. ## Flux Integration Flux can decrypt SOPS-encrypted files automatically. You configure a decryption secret once: ```yaml apiVersion: v1 kind: Secret metadata: name: sops-age namespace: flux-system stringData: age.agekey: | # created: 2025-01-15T10:30:00Z # public key: age1ql3z... AGE-SECRET-KEY-1ABCD... ``` Then reference it in Kustomization resources: ```yaml apiVersion: kustomize.toolkit.fluxcd.io/v1 kind: Kustomization metadata: name: infrastructure namespace: flux-system spec: interval: 10m path: ./gitops/infrastructure/production prune: true sourceRef: kind: GitRepository name: flux-system decryption: provider: sops secretRef: name: sops-age ``` Flux pulls encrypted manifests from git, decrypts them using the age key, and applies them to the cluster. You never handle plaintext secrets in CI pipelines. ## Practical Workflow **Encrypting a new secret**: ```bash # Create plaintext secret cat > database-credentials.yaml < database-credentials.enc.yaml # Commit the encrypted version git add database-credentials.enc.yaml git commit -m "Add database credentials for production" git push ``` **Editing an existing secret**: ```bash # SOPS decrypts, opens in editor, re-encrypts on save sops database-credentials.enc.yaml # Commit the changes git add database-credentials.enc.yaml git commit -m "Rotate database password" git push ``` **Rule**: Always encrypt before commit. We use pre-commit hooks to catch accidentally committed plaintext secrets. ## Key Rotation Rotating keys periodically limits the blast radius if a key is compromised: ```bash # Generate new age key age-keygen -o age-new.key # Re-encrypt all secrets with new key find gitops/ -name "*.enc.yaml" -exec sops rotate -i {} ; # Update .sops.yaml with new public key # Update Flux secret with new private key # Commit and deploy ``` We rotate keys annually and when team members with key access leave. ## What This Enables **Disaster recovery**: Entire infrastructure from git, including secrets. Fresh cluster? Clone repo, bootstrap Flux, done. **Audit trail**: Every secret change is a git commit. Who rotated the database password? Check git log. **Review process**: Secret changes go through pull requests. Someone reviews before production. **No credential sprawl**: No secrets in CI configuration, no secrets in chat history, no secrets on developer laptops (except the age key itself). SOPS with age keys isn't the only secrets solution, but it's the one that makes GitOps actually work for complete infrastructure. **Sources:** - [SOPS (Mozilla)](https://github.com/getsops/sops) - [age encryption](https://age-encryption.org/) - [Flux SOPS integration](https://fluxcd.io/flux/guides/mozilla-sops/) #sops #age #secrets #gitops #security --- ## Pull-Based GitOps: Why Flux Reduces Both Compute and Cognitive Load URL: https://clouds-of-europe.eu/content/strategy-transition/resource-efficiency-strategy/pullbased-gitops-why-flux-reduces-both-compute-and-cognitive Author: Jurg van Vliet Published: 2025-11-12 Category: Strategy & Transition Type: Resource Efficiency (Strategy) Tags: flux, gitops, operationalexcellence, security, sops ## Push vs Pull: A Trust Model Difference Most CI/CD systems push changes to production. Jenkins runs a job, authenticates to your cluster, and applies changes. GitHub Actions does the same. CircleCI, GitLab CI—all push-based. This creates a security boundary problem: your CI system needs production credentials. If CI is compromised, production is compromised. Flux inverts this model. The cluster pulls its own configuration from git. CI never touches production. The trust boundary shifts. ## How Flux Actually Works Flux runs inside your Kubernetes cluster. It watches your git repository. Every few minutes (configurable), it checks: does the cluster state match what's in git? If not, Flux reconciles the difference. Resources are created, updated, or deleted to match git. Manual changes get reverted. Git is the source of truth, enforced continuously. **What this means operationally**: - CI pipelines never need production credentials - Manual changes don't persist (drift is automatically corrected) - Every change is auditable (it's a git commit) - Rollback is `git revert` followed by automatic reconciliation ## The Credential Problem, Solved Before Flux, our GitHub Actions needed: - Kubernetes API credentials for the test cluster - Kubernetes API credentials for the production cluster - Secrets for each environment, stored in GitHub This is a lot of high-privilege credentials floating around. Every engineer with repo access could potentially see or use them. Every supply chain vulnerability in the CI pipeline was an exposure. With Flux: - GitHub Actions build containers, push to registry—that's it - No production credentials in CI - Flux (running in production) pulls the image and applies configuration - Credentials stay inside the cluster **The attack surface reduction is substantial.** Compromise CI, and you can poison the container registry—but you can't directly access production infrastructure. It's not perfect security (no such thing), but it's a meaningful improvement. ## SOPS: Encrypted Secrets in Git "You can't commit secrets to git" is conventional wisdom. SOPS (Secrets OPerationS) challenges that. SOPS encrypts secret values while leaving structure readable: ```yaml apiVersion: v1 kind: Secret metadata: name: database-credentials stringData: password: ENC[AES256_GCM,data:Zq7X...,iv:abc...,tag:def...,type:str] ``` The file is committed to git. The encrypted value is useless without the decryption key. Flux has the key (stored as a Kubernetes secret, never committed). During reconciliation, Flux decrypts automatically. **What you gain**: - Secret management follows the same workflow as everything else (git PR, review, merge) - Audit trail for secret changes (who changed what, when) - Disaster recovery includes secrets (reproduce the entire cluster from git + decryption key) **The tradeoff**: Key management is now your responsibility. We use age keys, rotated annually, backed up securely. It's overhead, but less than managing secrets across multiple systems. ## Self-Healing: The Cognitive Load Benefit Before GitOps, if someone manually changed a production setting, that change persisted. You'd discover it later—maybe during an incident—and wonder: who changed this? Why? Is it intentional or accidental? With Flux, manual changes get reverted within minutes. The cluster continuously reconciles toward git. This is self-healing infrastructure. **The psychological benefit is larger than you'd expect.** You can trust the cluster. You know what's running (look at git). You know manual changes won't stick. This reduces anxiety. When you're on call at 3 AM debugging an incident, you're not wondering if someone made an undocumented change. The running state matches git. That certainty reduces cognitive load meaningfully. ## Resource Efficiency: Fewer CI Runners Push-based CI means every deployment runs a CI job. That job needs compute: a VM or container, running for minutes, authenticating, applying changes, then terminating. Flux runs continuously, consuming ~100-200MB RAM and minimal CPU. One Flux controller handles dozens of applications. The compute cost is fixed and small. **Rough math**: Before Flux, our GitHub Actions deployment jobs consumed maybe 10-15 minutes of runner time per deployment. At ~20 deployments per day, that's 200-300 minutes of compute. Flux runs continuously but uses less total compute than 30 minutes of CI runners. This isn't massive savings, but it's directionally positive. Less compute means lower cost and lower carbon footprint. ## The Human Sustainability Angle Better sleep is a legitimate operational metric. With push-based deployments, you're responsible for every step. Build fails? You fix it. Deployment script breaks? You fix it. Credentials expire? You fix it at 2 AM when a deployment fails. With Flux, the system is self-reconciling. Most failures are transient—Flux retries automatically. Configuration errors are caught in pull requests (dry-run validation). The system tolerates temporary failures gracefully. **This changes on-call quality.** Issues still happen, but the system is more resilient by default. You're debugging application logic, not deployment mechanics. That's a better use of human attention. ## Getting Started with Flux If you want to try this model: ```bash # Bootstrap Flux in your cluster flux bootstrap github \ --owner=your-org \ --repository=your-repo \ --path=clusters/production \ --personal # This sets up Flux to watch your repo's clusters/production directory # Put Kubernetes manifests there, commit, push # Flux applies them automatically ``` Start simple: one application, standard Kubernetes manifests. Once you understand the pattern, add SOPS for secrets, then expand to more applications. The learning curve is real but manageable. The operational benefits—security, auditability, reduced cognitive load—compound over time. **Sources:** - [Flux CD documentation](https://fluxcd.io/) - [SOPS (Mozilla)](https://github.com/getsops/sops) - [GitOps principles](https://opengitops.dev/) #gitops #flux #sops #security #operationalexcellence --- ## Centralized Monitoring: One Pane of Glass, Better Sleep URL: https://clouds-of-europe.eu/content/strategy-transition/collaborative-infrastructure/centralized-monitoring-one-pane-of-glass-better-sleep Author: Jurg van Vliet Published: 2025-10-15 Category: Strategy & Transition Type: Collaborative Infrastructure Tags: centralizedinfrastructure, grafana, monitoring, observability, prometheus ## The Per-Environment Monitoring Problem Standard Kubernetes monitoring: every cluster runs Prometheus, Grafana, and associated infrastructure. This is recommended practice—monitoring should be reliable even when the application is failing. For a single cluster, this works fine. For multiple environments (development, test, production), it becomes repetitive: - Three Prometheus instances, each scraping their own cluster - Three Grafana instances, each with their own dashboards - Three sets of alerting rules, which inevitably drift out of sync - Three places to look during an incident Resource cost is real—Prometheus and Grafana aren't lightweight. But the cognitive cost is larger. During an incident, which Grafana are you looking at? Are the dashboards identical? Is the alert configuration the same? ## The Centralized Alternative We run a management cluster—a small Kubernetes cluster whose only job is monitoring and GitOps control for other clusters. Each workload cluster runs Prometheus Agent (not full Prometheus). The agent scrapes metrics and remote-writes them to Mimir in the management cluster. Grafana in the management cluster queries Mimir for metrics from all environments. **What this looks like**: ```yaml # In each workload cluster: Prometheus Agent apiVersion: v1 kind: ConfigMap metadata: name: prometheus-agent-config data: prometheus.yml: | remote_write: - url: https://mimir.mgmt.example.com/api/v1/push basic_auth: username: password_file: /etc/prometheus/secrets/password ``` Prometheus Agent is much lighter than full Prometheus—no local storage, no query engine, just scraping and forwarding. Memory footprint drops from 2-4GB to 200-500MB. ## Resource Savings Per cluster, we eliminated: - Full Prometheus: ~2-4GB RAM, 2+ CPU cores - Grafana: ~500MB RAM, 1 CPU core - Persistent storage for metrics: 50-100GB Across three clusters (dev, test, production), that's roughly: - 6-12GB RAM saved - 6-9 CPU cores freed - 150-300GB storage eliminated The management cluster runs Mimir (for metrics storage) and Grafana (for visualization), but that's shared across all environments. One Grafana instance serves dashboards for all clusters. **Total resource reduction: roughly 60-70%.** Not revolutionary, but meaningful at scale. ## The Real Win: Cognitive Simplicity During an incident at 3 AM, you're half-awake, stressed, and need answers fast. With centralized monitoring: - One URL to remember: monitoring.example.com - One dashboard showing all environments - One set of alerts in one place - Correlate across clusters easily (did test see this issue earlier?) Before centralization, I'd have three browser tabs open, trying to remember which Grafana showed production. It sounds trivial. At 3 AM, it's friction you don't need. **Cognitive load reduction is real.** One interface means less mental overhead. Unified dashboards mean consistent visualization. This improves incident response measurably. ## Security: mTLS for Remote Write Remote write means metrics leave the workload cluster and travel to the management cluster. This is cross-network traffic; it needs security. We use mTLS (mutual TLS) for Prometheus remote write: - Each Prometheus Agent has a client certificate - Mimir requires valid certificates - Certificates are issued by cert-manager from a cluster CA - Traffic is encrypted and authenticated **Configuration example**: ```yaml remote_write: - url: https://mimir.mgmt.example.com/api/v1/push tls_config: cert_file: /etc/prometheus/certs/tls.crt key_file: /etc/prometheus/certs/tls.key ca_file: /etc/prometheus/certs/ca.crt ``` This isn't just security theater. In regulated environments, encrypting and authenticating metrics traffic is a requirement. mTLS provides both with standard tooling. ## Tradeoffs: Single Point of Failure The obvious concern: if the management cluster fails, do you lose all monitoring? Yes and no. Prometheus Agent buffers metrics locally. If remote write fails (network issue, management cluster down), the agent queues metrics and retries. Once connectivity returns, metrics backfill automatically. **How long can it buffer?** Depends on memory and retention settings. We configure 1-2 hours of buffer. For most incidents, that's sufficient. **What if management cluster is down longer?** You still have application logs and Kubernetes events in the workload cluster. You can deploy Grafana temporarily to that cluster if needed. It's not ideal, but it's a fallback. The tradeoff is acceptable: simplified operations 99% of the time, with a known recovery path for the 1% edge case. ## Why This Matters for Small Teams Large organizations can afford dedicated monitoring specialists. Small teams need infrastructure that reduces overhead, not increases it. Centralized monitoring means: - One system to learn instead of three - One place to maintain dashboards and alerts - Lower resource costs (matters when you're cost-conscious) - Better incident response (matters when you're on call) The pattern scales down effectively. Even two clusters benefit from centralization. The cognitive simplification pays for itself. **Sources:** - [Prometheus Agent Mode](https://prometheus.io/docs/prometheus/latest/feature_flags/#prometheus-agent) - [Grafana Mimir](https://grafana.com/oss/mimir/) - [cert-manager documentation](https://cert-manager.io/docs/) #monitoring #prometheus #grafana #centralizedinfrastructure #observability --- ## What has OpenHW todo with tech sovereignty for embedded electronics URL: https://clouds-of-europe.eu/content/policy-sovereignty/open-source-standards/what-has-openhw-todo-with-tech-sovereignty-for-embedded-electronics Author: Jasper Geurtsen Published: 2025-09-28 Category: Policy & Sovereignty Type: Open Source & Standards **OpenHW Group** is directly tied to tech sovereignty for embedded electronics, because it’s one of the key organizations enabling open-source processor ecosystems that can replace or complement proprietary IP (like ARM or proprietary RISC cores). ---- Here’s how OpenHW fits in: 🔧 What is OpenHW Group? - A nonprofit industry consortium (founded 2019). - Builds open-source RISC-V processor cores and related IP for use in embedded systems. - Focused on industrial-grade quality—meaning verified, supported, and suitable for production (unlike some academic-only cores). - Example: the CORE-V family of RISC-V cores. ---- 🧩 Connection to Tech Sovereignty 1. Alternative to proprietary architectures - Most embedded electronics today rely on ARM Cortex-M or Cortex-A cores. - OpenHW provides open, royalty-free cores based on RISC-V, so companies (or countries) aren’t locked into ARM licensing or export controls. 2. Transparency & Trust - Open-source RTL (Verilog/VHDL) lets engineers and governments audit the hardware IP. - Important for security-sensitive applications (defense, industrial, healthcare). 3. Ecosystem Independence - By standardizing around open cores, companies can develop their own SoCs without relying on proprietary ARM toolchains or licensing regimes. - Reduces dependency on single vendors and strengthens sovereignty. 4. Community & Skills - OpenHW builds a shared ecosystem of verification, testbenches, and software support. - This makes it easier for local engineering teams (universities, SMEs, startups) to adopt sovereign hardware instead of being tied to global proprietary players. ---- 🚀 Practical Example Imagine you’re building an industrial controller or automotive ECU: - With ARM: you depend on ARM’s licensing, NDA-bound docs, and vendor roadmaps. - With OpenHW’s CORE-V cores: You can integrate an open RISC-V core into your chip design. You own the RTL, can verify it yourself, and even modify it. You can use open toolchains (LLVM, GCC, Zephyr RTOS). You’re no longer at risk if ARM raises royalties or geopolitical restrictions cut off access. ---- So: **OpenHW** is one of the enablers of embedded tech sovereignty, because it provides trusted, open, production-ready processor cores that free companies from dependence on proprietary, foreign-controlled IP. --- ## Cloud TCO: Competitive Pricing with Sovereignty Included URL: https://clouds-of-europe.eu/content/strategy-transition/business-cases/cloud-tco-competitive-pricing-with-sovereignty-included Author: Jurg van Vliet Published: 2025-09-20 Category: Strategy & Transition Type: Business Cases Tags: cloudpricing, europeancloud, finops, honestnumbers, tco ## The Pricing Claim You Won't See You won't see me claim "Scaleway is 50% cheaper than AWS." I've reviewed enough cost analyses to know that claim doesn't hold up. Here's the honest assessment: after year-one credits expire, European providers are **competitively priced** with hyperscalers for equivalent services. Sometimes slightly cheaper, sometimes slightly more expensive, usually within 10-20%. The value proposition isn't primarily cost. It's everything else that comes with the choice. ## Breaking Down Real Costs Let's look at actual pricing for a representative production workload: **Compute (Kubernetes nodes)**: - AWS EKS: t3.medium (2 vCPU, 4GB RAM) = ~€35/month × 3 nodes = €105/month - Scaleway Kapsule: DEV1-M (3 vCPU, 4GB RAM) = ~€30/month × 3 nodes = €90/month Scaleway is ~15% cheaper for compute. Not revolutionary. **Managed Kubernetes control plane**: - AWS EKS: €70/month per cluster - Scaleway Kapsule: Free This is a real difference—€70/month saved per cluster. For multiple environments (dev, test, prod), that adds up. **Object Storage (S3-compatible)**: - AWS S3: €0.023/GB/month (standard tier) - Scaleway Object Storage: €0.01/GB/month For 1TB storage: AWS €23/month, Scaleway €10/month. Scaleway wins on storage pricing. **Database (managed PostgreSQL)**: - AWS RDS: db.t3.medium = ~€90/month + storage - Scaleway Database: DB-DEV-M (similar spec) = ~€60/month + storage Roughly 30% cheaper for Scaleway on databases. **Network egress (the big one)**: - AWS: €0.09/GB for most egress - Scaleway: €0.01/GB after 75GB free per month This is where costs diverge significantly. For a typical application serving 5TB/month: - AWS: €450/month in egress alone - Scaleway: ~€50/month (after free tier) Egress pricing is genuinely different, and it matters for data-intensive workloads. ## Total Cost Comparison For a realistic production workload: - 3-node Kubernetes cluster - Managed PostgreSQL database - 1TB object storage - 5TB/month egress **AWS**: ~€700-800/month **Scaleway**: ~€300-400/month Scaleway comes out 40-50% cheaper in this scenario. But there are caveats. ## The Caveats You Need to Know **1. Year-one credits distort everything** AWS, Azure, and GCP offer generous startup credits. €100K in credits makes everything look free for the first year. Scaleway offers credits too, but usually smaller amounts. The fair comparison is **year three pricing**, after all credits expire. That's when you see real costs. **2. Proprietary services cost more** The comparison above uses commodity services: compute, storage, databases. If you use AWS Lambda, DynamoDB, SQS, or other proprietary services, costs change. Proprietary services are priced at premium margins. Equivalent functionality on standard infrastructure (Kubernetes, PostgreSQL, Redis) is usually 30-50% cheaper—but you're managing more yourself. **The tradeoff**: Convenience and features vs portability and cost. Neither answer is universally right. **3. Reserved capacity vs on-demand** Hyperscalers offer significant discounts (30-50%) for reserved instances. If you commit to 1-3 years, costs drop substantially. European providers typically offer smaller discounts for long-term commitments. Reserved instance pricing for AWS can be competitive with or cheaper than Scaleway on-demand. **But**: Reserved instances are lock-in. You've committed. Negotiating leverage decreases. ## The Optionality Value Here's what doesn't show up in spreadsheets: optionality. When renewal time comes and you're locked into AWS proprietary services with reserved instances, your negotiating position is weak. "Move to another provider" isn't credible—the migration cost is prohibitive. When you've built on Kubernetes with portable services and aren't locked into proprietary features, "we're considering OVHcloud" is a real conversation. Even if you don't switch, the credible option affects pricing. **How to value this**: The option to leave has value even if you never exercise it. Call it a 10-20% discount you didn't have to explicitly negotiate. ## What Actually Matters: The Pitch to Leadership When presenting to non-technical leadership, cost matters. But it's not the whole story. **The pitch we use**: "Scaleway is cost-competitive with AWS—roughly equivalent after credits expire, sometimes cheaper for our workloads. The core value isn't price savings. It's: 1. **European jurisdiction**: Data stays under European law, simplifying GDPR compliance 2. **Reduced vendor lock-in**: We maintain the ability to migrate, which improves our negotiating position 3. **Team capability**: We build deeper infrastructure knowledge, which serves us long-term 4. **Predictable costs**: Fixed-price capacity reduces surprise bills and budget uncertainty The 20% workload experiment costs roughly €X/month. If it works, we gain leverage and optionality. If it doesn't, we've learned something valuable for €X." Notice: cost is mentioned, but it's not the primary argument. That's honest positioning. ## When European Providers Are Actually Cheaper European providers have genuine cost advantages for: **High egress workloads**: Serving large files, video streaming, or data-intensive APIs. Hyperscaler egress costs dominate; European providers are 5-10x cheaper here. **Simple infrastructure**: If you're running VMs or Kubernetes without many managed services, European providers are straightforward and competitively priced. **Multi-environment setups**: Free Kubernetes control planes add up when you're running dev, test, staging, and production clusters. ## When Hyperscalers Make Sense Conversely, hyperscalers have advantages for: **Global reach**: If you need presence in 20+ regions worldwide, AWS/Azure have broader coverage. **Exotic managed services**: If your architecture depends on services like AWS Lambda, Step Functions, or DynamoDB, equivalents elsewhere require rearchitecting. **Enterprise support contracts**: Large organizations with negotiated enterprise agreements get pricing and support that changes the calculation. ## The Honest Bottom Line After year-one credits expire, Scaleway and similar European providers are **competitive but not dramatically cheaper** than hyperscalers. The value is: - European jurisdiction and compliance simplification - Reduced lock-in and improved negotiating position - Operational model that builds team capability - Predictable, understandable pricing These are real benefits. Overselling the cost savings undermines credibility. The value proposition is broader than price. **Sources:** - [Scaleway Pricing](https://www.scaleway.com/en/pricing/) - [AWS Pricing Calculator](https://calculator.aws/) #tco #cloudpricing #finops #europeancloud #honestnumbers --- ## Your First Migration: The Pioneer Project Strategy URL: https://clouds-of-europe.eu/content/strategy-transition/strategic-planning/your-first-migration-the-pioneer-project-strategy Author: Jurg van Vliet Published: 2025-09-10 Category: Strategy & Transition Type: Strategic Planning Tags: cloudstrategy, incrementalapproach, kubernetes, learningbyexecuting, migration ## Why Not "Lift and Shift Everything"? I've seen too many ambitious cloud migration plans: "We'll move all 47 applications to European infrastructure in Q2." By Q3, they're still on slide 12 of the migration plan, and leadership is losing patience. Wholesale migration is the wrong approach. Instead, pick one workload. Get it right. Learn. Iterate. ## Criteria for Your Pioneer Project The ideal first migration has four characteristics: **1. Stateless (or externalized state)** A stateless API, a web frontend, or a background worker. If the workload stores data locally, it's harder. If it uses an external database or object storage, that's fine—you're not migrating the data layer yet. **Why this matters**: Stateless workloads are disposable. If something goes wrong, you can switch back instantly. You're testing the deployment process without the complexity of data migration. **2. Already containerized** If it runs in Docker or on Kubernetes today, it can run on any Kubernetes. If it's not containerized, containerize it first—that's a separate project with its own learning curve. **Why this matters**: Containerization is provider-neutral. Once you're in a container, the underlying platform becomes largely irrelevant. This is the portability layer. **3. Low business risk** Development environments, internal tools, or non-customer-facing services are ideal. If something goes wrong—deployment fails, latency increases, service goes down—the blast radius is contained. **Why this matters**: You're learning. Learning involves mistakes. Make those mistakes where they won't affect customers or revenue. **4. Representative complexity** Don't pick a "hello world" static page. Pick something that exercises your actual infrastructure: multiple services, database connections, external integrations, monitoring. **Why this matters**: A trivial workload won't reveal real challenges. You need enough complexity to test your processes, find hidden dependencies, and build realistic confidence. ## Good First Migrations Examples I've seen work well: **Internal admin panel**: Used by your team, not customers. Connects to the same database your production apps use. Exercises authentication, API access, and deployment workflows. Low risk, representative complexity. **Development/staging environment**: Already separate from production. Acceptable downtime. Same stack as production, so you're testing the full workflow. Ideal learning ground. **Background job processor**: Pulls work from a queue, processes it, writes results. Stateless, retriable work. Can run in parallel with existing workers during transition. **Documentation or marketing site**: Low traffic, static or nearly static, but uses your real deployment pipelines. Good for testing CI/CD integration. ## Bad First Migrations **Your billing system**: State-heavy, zero-downtime requirement, catastrophic if wrong. Save this for after you've proven the pattern. **Legacy application with undocumented dependencies**: "We think it only talks to the database, but we're not sure." Figure out the dependencies first. **Anything with compliance unknowns**: If you don't know whether the workload has GDPR, PCI, or regulatory requirements, research that before starting. **Your highest-traffic service**: You want to learn on something forgiving, not something where 0.1% error rate means angry customers. ## What You Learn from the Pioneer A successful first migration teaches you: **Your actual portability**: Where are the hidden cloud provider dependencies? AWS-specific SDKs, proprietary services, hard-coded regions—these reveal themselves during migration. **Team capability**: Who on the team can manage Kubernetes? Who understands networking? Where are the knowledge gaps that need filling? **Deployment patterns**: Does your CI/CD work with a new provider? Are credentials managed properly? Does the deployment automation work or is it tied to specific environments? **Operational readiness**: Can you monitor, debug, and respond to incidents in the new environment? Is observability set up correctly? **Cost reality**: What does this actually cost to run? Are the estimates accurate? Where are the unexpected charges? ## After Success: The 20% Experiment Once your pioneer project is stable, you have proof. You've demonstrated: - The provider works for your workloads - Your team can operate the infrastructure - Costs are acceptable - Performance meets requirements Now you can make a realistic decision about the next 20%. Not everything needs to move—maybe only certain workloads benefit from European infrastructure. But you're making that decision from experience, not speculation. The pioneer project isn't about migrating everything. It's about building capability and proving viability. Do it well, and the rest becomes tractable. #migration #kubernetes #cloudstrategy #incrementalapproach #learningbyexecuting --- ## Building Engineering Capability: The 3-6 Person Model URL: https://clouds-of-europe.eu/content/strategy-transition/leadership-guides/building-engineering-capability-the-36-person-model Author: Jurg van Vliet Published: 2025-08-25 Category: Strategy & Transition Type: Leadership Guides Tags: capability, engineering, leadership, sustainability, teambuilding ## The Real Constraint After years of infrastructure transitions, I'm convinced the hard part isn't technical. It's organizational. "Nobody gets fired for buying AWS" is the modern version of the IBM adage. It's true—there's safety in choosing the dominant vendor. If something goes wrong, you made the safe choice. Choosing a different path means accepting responsibility for outcomes you could have outsourced. That's a real consideration, not something to dismiss. ## The "Too Small to Run Kubernetes" Myth The most persistent objection: "You need a huge DevOps team to run Kubernetes." This was true in 2015. Kubernetes was immature, operationally complex, poorly documented. Running it required specialists. It's not true in 2025. **What changed:** - **Managed Kubernetes** eliminates control plane operations (no etcd management, no API server upgrades) - **GitOps** (Flux, ArgoCD) reduces manual operations and prevents drift - **Mature ecosystem** provides patterns, tools, community support (Stack Overflow has answers) - **AI assistance** accelerates learning and reduces toil The operational burden has dropped dramatically. The knowledge required is more accessible. The tooling is better. ## The 3-6 Person Model This is our actual operational model. Not hypothetical—this is what works for us. **Why 3 people minimum?** On-call rotation. If you want sustainable 24/7 coverage with reasonable quality of life, you need at least 3 people rotating. **With 3 people:** - Each person: 1 week on-call, 5 weeks off (rotation every 6 weeks) - Coverage during vacation (2 people can cover while 1 is away) - Knowledge sharing (rotate, learn from each other) - Resilience to attrition (losing 1 person doesn't break coverage) **With 2 people:** - Each person: 1 week on, 1 week off (unsustainable) - No vacation coverage (always on-call when not on vacation) - No resilience (1 person leaving breaks everything) - Burnout guaranteed **Why 6 people maximum?** Communication overhead. Beyond 6 people, coordination costs exceed productivity gains. Small teams move faster, make decisions easier, maintain shared context naturally. Brooks' Law applies: adding people to a late project makes it later. Keep the team small enough to maintain high communication bandwidth. ## What 3-6 People Can Actually Do With modern tooling, this team size can operate: **Infrastructure:** - Multi-cluster Kubernetes (management, test, production) - GitOps deployment pipelines (Flux reconciling from Git) - Centralized monitoring and observability (Prometheus, Grafana, Loki) - Secrets management (SOPS encryption, cert-manager certificates) **Operations:** - On-call rotation (1 week every 2 months per person) - Incident response (runbooks, clear escalation) - System maintenance (updates, capacity planning) - Security patching (responded to CVE in 3 hours) **Development:** - Feature implementation (new endpoints, UI improvements) - Test coverage (264 API tests, comprehensive E2E suite) - Documentation (architecture decisions, operational guides) - Refactoring and technical debt management This isn't theoretical. This is what we actually do. ## The Key: Leverage Small teams succeed through leverage—making each person's effort go further. **1. Managed services (selective use)** Don't run Kubernetes control planes. Don't manage PostgreSQL at the disk level. Don't operate email servers. Use managed offerings that reduce operational burden **while maintaining portability**: - Managed Kubernetes (Scaleway Kapsule, OVHcloud Managed K8s) - Managed PostgreSQL (CloudNativePG operator or provider-managed) - Transactional email (Scaleway TEM, SendGrid) Avoid managed services that create lock-in: - Proprietary serverless (Lambda, Cloud Functions) - Proprietary databases (DynamoDB, CosmosDB) - Proprietary queues (SQS, Cloud Pub/Sub) **2. GitOps automation** Flux reconciles automatically. Clusters pull their own configuration. Manual deployment eliminated. The system maintains itself within defined boundaries: - Drift is prevented, not detected and manually fixed - Changes are auditable (Git history shows everything) - Rollback is clean (`git revert` + wait for reconciliation) This reduces operational toil significantly. No "deploy to production" procedures. Just commit to Git. **3. Comprehensive monitoring** Good observability means you understand problems faster. Our centralized monitoring: - Single pane of glass (one Grafana, all environments visible) - Complete data (no sampling, year-plus retention) - Fast debugging (correlate metrics and logs in seconds) During incidents, time to understanding matters. Good monitoring reduces MTTR (mean time to recovery) dramatically. **4. AI assistance** Claude helps us understand complex systems faster. It's not replacing engineers—it's multiplying their effectiveness. Real examples: - Debugging multi-cluster networking issues - Writing OpenTofu modules that follow best practices - Understanding certificate chain validation failures - Transforming architecture discussions into documentation **5. Focus on high-value work** Automate toil. The team should work on: - Architecture and design decisions - Feature development - System improvements - Knowledge building Not on: - Manual deployments (automated via GitOps) - Configuration drift fixes (prevented by Flux) - Routine monitoring (automated alerts) - Repetitive operations (scripted or eliminated) ## The Sustainable On-Call Model On-call rotation must be sustainable. Burnout prevents long-term success. Here's our model: **Rotation schedule:** - **1 week on-call, 5-7 weeks off** (with 3-person rotation) - Predictable schedule (planned months in advance) - No "always on-call" culture - Clear handoff procedures **Escalation paths:** - Primary on-call person handles initial response - Secondary on-call for escalation (clear criteria for when to escalate) - Full team escalation only for critical incidents affecting revenue **Runbooks for common issues:** - Clear procedures reduce decision-making during stress - Known problems have documented solutions - Links to relevant dashboards and log queries - Escalation criteria defined explicitly **Blameless postmortems:** - Learning, not punishment - What went wrong? What went right? - How can we prevent recurrence? - What should we improve? This prevents burnout. Engineers can plan their lives, take vacations, have weekends—not be constantly anxious about alerts. ## Skills vs Headcount Small teams need **breadth of skills**, not narrow specialization. **T-shaped skill requirements:** - **Deep expertise in one area** (frontend, backend, infrastructure, databases) - **Working knowledge across stack** (can debug, understand tradeoffs, collaborate) This isn't about "10x engineers"—it's about **team members who can work across the stack when needed**. Practical example: - Frontend engineer can debug API issues (understands HTTP, REST, auth) - Backend engineer can troubleshoot Kubernetes pods (understands containers, networking) - Infrastructure engineer can review application code (understands application requirements) This breadth enables small teams to move fast without constant handoffs. ## What Doesn't Work Don't try to build everything. Small teams fail when they: **1. Run their own Kubernetes control planes** Use managed Kubernetes. Control plane operations (etcd management, API server upgrades, scheduler tuning) are toil that doesn't differentiate your business. **2. Build custom CI/CD from scratch** Use GitHub Actions, GitLab CI, or similar. Good CI/CD exists. Don't rebuild it. **3. Implement every feature themselves** Use mature open source tools (Prometheus, Grafana, PostgreSQL, cert-manager). Stand on shoulders of giants. **4. Operate 24/7 with 2 people** Burnout guaranteed. You need minimum 3 people for sustainable on-call rotation. **5. Optimize prematurely** Start with simple architecture. Add complexity only when you have evidence it's needed. "Might need to scale to millions" isn't evidence. ## Recruiting for Small Teams Curious engineers want interesting problems. **"We configure managed services"** is less compelling than **"We build and operate our own infrastructure."** The engineers we want to attract: - Care about understanding how systems actually work (not just clicking buttons) - Value building over assembling (create, don't just compose) - Think about long-term consequences (architecture that lasts) - Are motivated by mission (sovereignty, privacy, sustainability matter) Running on European infrastructure, building on open standards, maintaining portability—these aren't just operational choices. They're **signals about engineering culture**. Engineers who care about these things often care about: - Code quality and craftsmanship - System understanding and debugging skill - Long-term thinking and sustainability - Impact beyond just shipping features This is the talent you want. The mission helps attract them. ## The Honest Tradeoffs Small team benefits: - Fast decision-making (no lengthy approval processes) - High context (everyone knows everything) - Direct communication (talk, don't email) - Shared ownership (everyone responsible for everything) Small team costs: - Limited specialization (everyone wears multiple hats) - Vacation coverage challenge (3-person minimum for reason) - Knowledge concentration risk (bus factor) - Hiring pressure (each hire has huge impact) For European cloud independence, small teams are often ideal: **you can move fast, make bold architectural choices, and build capability without bureaucracy**. ## Making It Work **Clear responsibilities:** Even in small teams, someone needs to own each area. Ownership doesn't mean "only person who works on it"—it means "person responsible for ensuring it works." **Regular knowledge sharing:** Weekly technical discussions, architecture reviews, postmortems. Keep everyone's context current. **Documentation discipline:** Small teams are tempted to skip documentation ("we all know this"). Don't. You'll forget. New people will join. Document as you build. **Sustainable pace:** Marathon, not sprint. Protect against burnout. Enforce reasonable on-call rotations. Take vacations. **Key Points:** - 3-6 engineers can run production Kubernetes infrastructure - Managed services + GitOps + AI assistance = force multipliers - Sustainable on-call: 1 week every 2 months with 3-person rotation - T-shaped skills: breadth across stack, depth in one area - Small teams move fast with right tooling and patterns - Mission matters: sovereignty resonates with purpose-driven engineers #leadership #engineering #teambuilding #capability #sustainability --- ## Real Stories: Six Months to Production with AI Assistance URL: https://clouds-of-europe.eu/content/strategy-transition/success-stories/real-stories-six-months-to-production-with-ai-assistance Author: Jurg van Vliet Published: 2025-08-15 Category: Strategy & Transition Type: Success Stories Tags: aiassisted, europeancloud, gitops, kubernetes, successstory ## The Timeline June 3, 2025: Initial commit. "Reclaiming our digital independence." December 12, 2025: Production-ready platform with: - Multi-cluster Kubernetes (management, test, production) - Complete GitOps with Flux v2 and SOPS-encrypted secrets - Centralized monitoring (Prometheus, Grafana, Loki, Mimir) - 264 API tests with 100% endpoint coverage (58 endpoints) - Modern E2E test suite with Playwright - Production deployment on European infrastructure (Scaleway) **Six months.** Small team. Running a company. Built on European values throughout. ## Context That Matters This isn't a case study from a huge organization with unlimited resources. This is proof that **small teams can build European cloud independence fast** with modern tooling. **Team size:** 3-6 people (not dedicated full-time to this project) **Other constraints:** Running Aknostic (our company), client work, other responsibilities. This project was built in parallel with running a business, not instead of it. **AI assistance:** Claude as learning accelerator and pair programmer. This was explicitly an experiment: can small teams use AI tooling to build production systems faster? **Building on European values:** Sovereignty by design, not retrofit. Every architectural decision considered: Where will data live? What laws apply? Which company benefits? ## What Made This Possible **1. Kubernetes as Foundation** We didn't build abstractions from scratch. We built on Kubernetes—a proven platform with massive ecosystem support. Managed Kubernetes (Scaleway Kapsule) meant we focused on applications, not control plane maintenance. Free control plane, pay only for worker nodes. This is the modern pattern: **use managed services that reduce operational burden while maintaining portability**. **2. GitOps from Day One** Everything in Git from the beginning: - Infrastructure definitions (OpenTofu) - Kubernetes manifests (deployments, services, Gateway API) - Configuration (Helm values, Kustomize overlays) - Secrets (SOPS-encrypted) - Documentation (architecture decisions, runbooks) This created natural guardrails: all changes reviewable in PRs, all deployments auditable in Git history, all configuration reproducible from repository. **3. AI as Learning Accelerator** Claude didn't write all our code. It helped us **understand complex systems faster**. Examples: - "Why isn't my Gateway attaching to this HTTPRoute?" → ReferenceGrants explained - "How do I configure Flux to decrypt SOPS secrets?" → Complete config with explanation - "Debug this cert-manager certificate chain issue" → Probable cause identified quickly AI as pair programmer and documentation assistant. This made complex technologies (Kubernetes networking, GitOps reconciliation, mTLS) learnable in days instead of weeks. **4. Modern Testing Approach** We invested in testing from the start: - **264 black-box API tests:** Every endpoint covered, pure HTTP testing - **E2E tests with Playwright:** User journey verification - **Independent test data:** Tests create own data with unique identifiers This isn't overhead—it's infrastructure for velocity. With comprehensive tests, we refactor confidently. Changes that break functionality get caught before deployment. **5. European Providers Ready for Production** In 2020, this might have been harder. European providers were catching up. Managed services were immature. Multi-AZ wasn't widely available. By 2025, European providers offer production-ready platforms: - Scaleway: Managed Kubernetes, object storage, DNS API, transactional email, container registry - Multiple regions, each with 3+ availability zones - Pricing competitive with hyperscalers - Support responsive and knowledgeable The infrastructure gap has closed. European independence is now practically achievable, not aspirational. ## The AI Software Engineering Experiment This project was also an experiment: **Can small teams use AI tooling to build production systems faster?** After six months, the answer: **Yes, with important caveats.** **What AI accelerated:** **Understanding unfamiliar systems:** Kubernetes networking, Flux reconciliation loops, cert-manager certificate chains, Gateway API ReferenceGrants—these are complex with subtle interactions. Traditional learning: Read docs, search Stack Overflow, trial-and-error, eventually understand (days to weeks). With Claude: Ask specific questions, get explained with context from our codebase, understand in minutes to hours. **Writing boilerplate:** OpenTofu modules follow patterns. Kubernetes manifests have structure. Helm values have conventions. AI excels at generating syntactically correct YAML/HCL from requirements. Review and adjust, but starting point is solid. **Debugging complex issues:** Multi-cluster connectivity not working? Certificate chain invalid? DNS resolution failing? Paste error logs, get probable cause suggestions, faster path to solution. Not magic—still requires verification and understanding—but accelerated. **Documentation transformation:** Architecture decisions made in conversations need to become structured documentation. AI helps: raw notes → ADRs, scattered thoughts → coherent guides. **What AI didn't replace:** **Architectural decisions:** Should we use centralized or per-environment monitoring? Gateway API or stick with Ingress? SOPS or external secret store? These require understanding tradeoffs, organizational context, resource constraints. AI explains options and tradeoffs, but humans must decide. **Domain expertise:** European sovereignty requirements, GDPR implications, regulatory landscape, business model considerations—these require specific knowledge. AI can explain regulations but can't substitute for compliance expertise or business judgment. **Quality judgment:** Is this code maintainable? Will this architecture scale? Is this abstraction premature? Will on-call engineers understand this at 3am? These require experience and operational wisdom. AI suggests approaches; humans judge appropriateness. **Operational experience:** What fails at 3am? What's actually difficult to debug? What causes alert fatigue? These come from running systems, not generating code. ## The Honest Assessment **Would we have built this without AI?** Yes, but slower. Conservative estimate: 9-12 months instead of 6. **Could a single person do this with AI?** Probably not sustainably. On-call burden alone requires multiple people. AI doesn't solve operational load or 24/7 coverage needs. **Is AI necessary for European cloud independence?** No. But it's a significant accelerator for small teams learning complex technologies. ## What We'd Tell Someone Starting Today **1. Start with Kubernetes** It's the portability layer that makes independence possible. Don't build custom orchestration. **2. Use managed services thoughtfully** Managed Kubernetes: yes (reduces operational burden). Managed everything: no (creates vendor lock-in). Find the balance. **3. Adopt GitOps immediately** Flux creates guardrails that prevent drift. Everything in Git means everything is auditable and reproducible. **4. Test from the beginning** Tests are infrastructure for confident changes. 264 API tests took time to write but enable fast, confident refactoring. **5. Leverage AI tooling appropriately** AI as learning accelerator and boilerplate generator: excellent. AI replacing engineering judgment: dangerous. Review everything critically. **6. Build on European infrastructure** The providers are ready. The gap has closed. You're not sacrificing quality for sovereignty—you're getting both. ## The Realistic Timeline Six months included: - Learning Kubernetes, Gateway API, Flux, SOPS - Architectural mistakes and refactoring - Running a business in parallel - Building comprehensive test coverage - Documentation and runbooks If you're focused full-time with existing Kubernetes knowledge: probably 3-4 months. If you're learning as you go (like we were): 6-8 months is realistic. If you're doing this part-time while running a business: 6-12 months. Set realistic expectations. Speed matters less than getting it right. ## Proof Points After six months, we have: **Technical maturity:** - Multi-cluster GitOps in production - Centralized observability with year-plus retention - Rapid security response (CVE patched in 3 hours) - Zero-downtime deployments via Kubernetes rolling updates **Operational sustainability:** - On-call rotation: 1 week every 2 months (3-person team) - Incident response: clear runbooks, good monitoring - Maintenance overhead: manageable for small team **Business viability:** - Competitive pricing vs. hyperscalers - European jurisdiction simplifies compliance - Platform proves sovereignty is practical, not just aspirational This is real. It works. Small teams can do this. **Sources:** - Project git history (2,143 commits June-December 2025) - [Clouds of Europe GitHub Repository](https://github.com/aknostic/clouds-of-europe) #successstory #kubernetes #gitops #europeancloud #aiassisted --- ## Strategic Diversification: Why Optionality Has Business Value URL: https://clouds-of-europe.eu/content/policy-sovereignty/corporate-responsibility/strategic-diversification-why-optionality-has-business-value Author: Jurg van Vliet Published: 2025-08-05 Category: Policy & Sovereignty Type: Corporate Responsibility Tags: businesscase, cloudsofeurope, cloudstrategy, diversification, multicloud ## The Business Model Behind "Free Tier" Hyperscalers offer generous startup credits and free tiers. AWS gives $100K+ in credits to promising startups. Google Cloud offers $300K for some programs. This isn't charity—it's customer acquisition with sophisticated economics. The model works because: once you've built on proprietary services (Lambda, DynamoDB, Cloud Functions), switching costs create retention. The credits get you started. Lock-in keeps you paying. I've seen this play out repeatedly: a startup uses $100K in AWS credits, builds on Lambda and DynamoDB for velocity, then finds themselves locked in when credits end. The monthly bill comes as a shock—€5K, €10K, €20K per month—but migration would cost more than staying. There's nothing wrong with understanding this model. Problems arise when you don't factor it into your planning. ## What Diversification Actually Looks Like I'm not suggesting you migrate everything to European providers tomorrow. That's not practical, and for many workloads it might not make sense. What I am suggesting is strategic diversification. **The 20% experiment:** Pick workloads that could run elsewhere without significant friction. Development environments, internal tools, or new projects are good candidates. Run them on a European provider—OVHcloud, Scaleway, or Hetzner. Learn what works and what doesn't. **What you gain:** - Your team builds multi-cloud expertise (skills that transfer across providers) - Your architecture becomes more portable (you'll find and fix hidden dependencies) - Your negotiating position improves (you can credibly discuss alternatives) - You have a tested fallback if you ever need one (proved, not theoretical) **The cost:** Some additional operational complexity (managing multiple providers). Some learning curve (different APIs, consoles, quirks). Usually worth it for the strategic value gained. ## The Honest Pricing Story Let's talk realistically about cost. Because if the business case requires claiming "European providers are 50% cheaper," we're on shaky ground. **The truth: European providers are competitively priced, not cheaper.** After year-one credits expire and you compare apples-to-apples (similar instance types, equivalent storage, comparable services), costs are roughly similar. Sometimes European providers are slightly cheaper, sometimes slightly more expensive, usually within 10-20% depending on workload mix and region. **What creates the "cheaper" perception:** **Egress costs:** AWS charges €0.09/GB for data transfer out. Scaleway charges €0.01/GB (plus 75GB free monthly). For egress-heavy workloads, this is significantly cheaper. But for compute-heavy workloads with minimal egress, the difference is small. **Proprietary service premium:** Lambda, DynamoDB, SQS—these services are convenient and charge premium margins. Equivalent functionality on standard infrastructure (Kubernetes, PostgreSQL, RabbitMQ) often costs 30-50% less. But it requires more operational expertise. **Contract leverage after year one:** This is the hidden cost. When credits expire and you're locked into proprietary services, renewal negotiations favor the vendor. I've seen increases of 20-40% at renewal for organizations that couldn't credibly threaten to migrate. ## The Optionality Value Some value doesn't appear in cost comparison spreadsheets. **Scenario: Contract renewal with AWS** Current spend: €10K/month. Credits expired. Fully built on AWS-specific services (Lambda, DynamoDB, CloudFormation, proprietary APIs everywhere). **Without alternatives:** - Weak negotiating position (vendor knows switching would be extremely costly) - Renewal pricing: €12-13K/month (+20-30% increase) - You accept because migration would cost €100K+ in engineering time - This repeats every renewal **With maintained portability:** - Strong negotiating position (built on Kubernetes, standard APIs, could migrate to Scaleway/OVHcloud) - Renewal discussion: "We're evaluating European alternatives" - This is credible, not bluff (you've maintained portable architecture) - Renewal pricing: €10K/month (flat) or minimal increase - Avoided €2-3K/month × 12 months × years = substantial savings **ROI of portability:** The small architectural overhead of maintaining portability: - Use Kubernetes instead of ECS/proprietary orchestration - Use PostgreSQL instead of DynamoDB - Use standard APIs instead of vendor-specific - Avoid deep integration with proprietary services This overhead pays for itself through negotiating leverage. The vendor knows you can leave. This changes pricing conversations. ## What We Actually Tell Leadership "European cloud providers are **competitively priced**. We're not promising 50% cost savings—that's overselling. What we're offering: 1. **Comparable pricing** (within 10-20% depending on workload, often cheaper for egress-heavy applications) 2. **European jurisdiction** (data subject to European law, simpler GDPR compliance) 3. **Negotiating leverage** (credible alternatives prevent vendor lock-in pricing at renewal) 4. **Strategic independence** (control our infrastructure stack, respond quickly to security issues) 5. **Innovation investment in Europe** (revenues to European companies, not US parents) The value proposition isn't primarily cost. It's **strategic**. We pay roughly the same but gain independence, optionality, and sovereignty." ## Making Infrastructure Choices Visible Every infrastructure decision has implications beyond the spec sheet. The question isn't whether to think about them—it's whether to think about them consciously or by default. When choosing providers, consider: - **Where is the provider incorporated?** What laws apply to your data? - **Where do profits flow?** European company or subsidiary of US parent? - **What's the energy source?** Actual renewable, nuclear, or fossil fuel? - **What's the long-term relationship?** Partnership or lock-in? These aren't abstract concerns. They affect compliance, costs, sovereignty, and your options down the road. **Sources:** - [Scaleway vs AWS Comparison](https://www.scaleway.com/en/scaleway-vs-aws/) - [Scaleway Pricing](https://www.scaleway.com/en/pricing/) - [AWS S3 Pricing](https://aws.amazon.com/s3/pricing/) #cloudstrategy #diversification #multicloud #businesscase #cloudsofeurope --- ## Cloud Sustainability: What Actually Matters and What's Marketing URL: https://clouds-of-europe.eu/content/policy-sovereignty/resource-efficiency-policy/cloud-sustainability-what-actually-matters-and-whats-marketi Author: Jurg van Vliet Published: 2025-07-28 Category: Policy & Sovereignty Type: Resource Efficiency (Policy) Tags: efficiency, europeancloud, greencloud, infrastructure, sustainability ## The Physical Reality "The Cloud" sounds weightless. In reality, it's concrete, copper, silicon, and electricity—lots of electricity. The International Energy Agency estimates data centers consumed about 460 terawatt-hours (TWh) in 2024, roughly 1.5% of global electricity consumption. This has grown at 12% annually over the last five years. Projections suggest this could double to reach 945 TWh by 2030—nearly 3% of global electricity demand. This isn't inherently bad. Digital services enable efficiencies elsewhere. Video calls prevent flights. Cloud computing reduces on-premises server sprawl. But it does mean our infrastructure choices have real environmental consequences. ## Evaluating Sustainability Claims Most hyperscalers advertise sustainability initiatives. Here's what the claims actually mean: **Renewable Energy Certificates (RECs):** The data center runs on grid power (which may include fossil fuels). The company purchases certificates equivalent to renewable energy generated elsewhere. This is accounting, not actual renewable power. **Power Purchase Agreements (PPAs):** The company contracts for renewable energy to be generated. Better than RECs—it funds new renewable capacity—but the data center still typically runs on grid power. **Direct renewable power:** The data center actually runs on renewable electricity. This is what Nordic facilities often achieve with hydroelectric and geothermal power. **Carbon neutral:** Usually means offsets, not actual zero emissions. Read the methodology. Planting trees to offset emissions is better than nothing, but worse than not emitting. ## What European Geography Offers Europe has genuine advantages in data center sustainability—advantages rooted in geography and infrastructure, not just accounting. **Nordic Countries: Actual Renewable Electricity** **Iceland:** - Geothermal and hydro power (100% renewable grid) - Natural cooling (cold climate reduces cooling energy) - No accounting tricks—data centers literally run on geothermal electricity **Sweden and Finland:** - Abundant hydro power - Cold climate (free cooling most of the year) - Strong grid renewable mix **Heat Recovery:** Stockholm provides a remarkable example of circular economy thinking. Data centers produce waste heat. Instead of discarding it, Stockholm's district heating network captures and distributes it. By 2022, Stockholm data centers provided enough waste heat to warm 30,000 apartments annually. Stockholm Exergi (the energy company) expects data centers to provide 10% of the city's heating needs. Close to 90% of Stockholm's buildings connect to this district heating network. Example: DigiPlex's Stockholm data center generates enough waste heat to warm 10,000 households. What was waste becomes value. This isn't greenwashing—it's actual resource efficiency built into urban infrastructure. **France: The Nuclear Question** Scaleway's France region (Paris) runs primarily on the French electrical grid, which is approximately 70% nuclear power. Nuclear power isn't ideal. Waste management requires centuries of careful oversight. Decommissioning costs are enormous. Accident risk, while low, carries catastrophic potential. But here's the pragmatic European position: **emissions cause more immediate harm than nuclear waste**. Climate change is happening now. Every ton of CO₂ emitted contributes to warming that affects billions of people. Nuclear waste is containable and localized—though requiring long-term management, it doesn't contribute to climate crisis. If the choice is between: 1. Fossil fuel emissions contributing to climate breakdown 2. Nuclear power with waste management challenges Nuclear is the more responsible short-term choice while we build out renewables. This doesn't make nuclear "good"—it makes it pragmatic given current reality. France has committed to nuclear as bridging technology until renewable capacity can fully replace fossil fuels. For data center operators: choosing French regions means choosing low-carbon electricity (not renewable, but low-carbon). This is better than fossil-powered alternatives. ## Practical Steps for Sustainability **1. Know your footprint** Ask your provider for actual energy source data, not just marketing claims: - What percentage of grid power is renewable? - What percentage is nuclear? - What percentage is fossil fuel? - Is this actual power source or accounting (RECs)? Reputable providers publish this data. If they won't share it, be skeptical. **2. Consider region selection** If latency permits, choose regions with cleaner grids: - Iceland: ~100% renewable (geothermal + hydro) - Sweden/Finland: ~60-70% renewable (hydro + wind) - France: ~70% nuclear (low-carbon, not renewable) - Germany: ~50% renewable (improving annually) Electricitymaps.com shows real-time grid carbon intensity. For batch workloads or development environments where latency is flexible, choose greener regions. **3. Right-size your workloads** A container requesting 2GB memory that uses 200MB wastes compute. This is money and energy. Measure actual usage with Prometheus: ```promql container_memory_usage_bytes / kube_pod_container_resource_requests{resource="memory"} ``` If you are over-provisioned, and your application allows it, set memory requests to actual usage plus 20-50% headroom. For CPU usage it is easier, as most applications handle this gracefully. **4. Question necessity** Do you need real-time processing, or would batch be acceptable? Batch jobs can run during high-renewable periods (windy days, sunny afternoons). Does this data need to live forever, or can you implement retention policies? Storage costs money and energy. Delete what you don't need. ## Efficiency Is Good Engineering Sustainability concerns align with good engineering practices: **Efficient code uses less compute:** - Faster response = less CPU time per request - Lower memory usage = more workloads per node - Optimized queries = less database CPU **Right-sized resources reduce waste:** - Accurate requests = better pod scheduling - Appropriate limits = prevent runaway consumption - Regular review = adapt to changing patterns **Fewer network hops mean lower latency and lower energy:** - Collocate related services - Cache aggressively - Minimize cross-region traffic You don't have to choose between performance and sustainability. Usually, they point the same direction. The code that runs fast also uses less energy. ## European Advantages Are Real Europe's geographic diversity provides genuine sustainability options: **Nordic regions:** Actual renewable electricity powering actual data centers. Stockholm turning waste heat into apartment heating. These aren't accounting tricks—they're infrastructure advantages. **France:** Low-carbon grid via nuclear. Not perfect, but pragmatic given climate urgency. **Germany:** Rapidly growing renewable capacity. Wind and solar deployment accelerating. When you choose European infrastructure, you can choose regions with genuinely cleaner electricity. This matters. **Sources:** - [IEA: Data centres & networks](https://www.iea.org/energy-system/buildings/data-centres-and-data-transmission-networks) - [Stockholm District Heating from Data Centers](https://eu-mayors.ec.europa.eu/en/Stockholm-Heat-recovery-from-data-centres) - [DigiPlex Stockholm: 10,000 Households Heated](https://datacentre.solutions/news/53792/digiplex-data-centre-to-heat-10000-stockholm-households) - [Carbon Brief: AI and Data Centre Energy Use](https://www.carbonbrief.org/ai-five-charts-that-put-data-centre-energy-use-and-emissions-into-context/) #sustainability #greencloud #efficiency #europeancloud #infrastructure --- ## Evaluating European Cloud Providers: Why Scaleway Shows the Way URL: https://clouds-of-europe.eu/content/policy-sovereignty/provider-ecosystem/evaluating-european-cloud-providers-why-scaleway-shows-the-w Author: Jurg van Vliet Published: 2025-07-22 Category: Policy & Sovereignty Type: Provider Ecosystem Tags: europeancloud, infrastructure, kubernetes, providerselection, scaleway ## What "Production-Ready" Actually Means There are many European hosting providers. Not all of them are cloud platforms suitable for production workloads. The difference matters. A VPS provider gives you virtual servers. A **cloud platform** provides an ecosystem: APIs for automation, managed services that reduce operational burden, and the resilience architecture to run without babysitting. For our workloads, we needed the latter. ## Our Evaluation Criteria Before evaluating providers, we defined what we actually needed: **1. Managed Kubernetes (non-negotiable)** We're a small team; we can't spend time managing control planes. We need managed Kubernetes that handles upgrades, scaling, and integration with the provider's networking and storage. Managed means: the provider operates the control plane (API server, scheduler, controller manager). We operate only worker nodes and workloads. **2. Multi-Availability Zone (production requirement)** A single data center is a single point of failure. For production, we require at least **three availability zones within a region**, allowing us to survive facility-level failures. This is table stakes for high availability. Without multi-AZ, you're accepting single-facility risk. **3. Essential Managed Services** Kubernetes alone isn't sufficient. You need primitives: - **S3-compatible Object Storage**: Backups, static assets, data storage - **Container Registry**: Local image storage (faster pulls, lower egress costs) - **Managed DNS with API**: Cert-manager integration for automatic TLS - **Transactional Email**: Because running your own mail server is avoidable pain These services need to work together. Container registry in Paris, object storage in Amsterdam, DNS API-driven—this creates the platform that enables independence. ## Why Scaleway Shows the Way Among European providers, Scaleway provides the most complete stack: **Managed Kubernetes (Kapsule):** - Multiple versions supported (1.31, 1.32, 1.33) - Automatic control plane upgrades - Integration with Scaleway networking, storage, DNS - Free control plane (pay only for worker nodes) - Multi-AZ within single region **Geographic Coverage:** - **Paris region**: 3 availability zones (PAR-1, PAR-2, PAR-3) - **Amsterdam region**: 3 availability zones - **Warsaw region**: 3 availability zones Each region independently provides multi-AZ deployment capability. This enables true high availability. **Object Storage (S3-compatible):** - Available across all regions - S3 API compatibility (works with standard tools) - Multi-AZ replication within region - Cross-region replication available - Pricing: €0.01/GB/month storage **Container Registry:** - Integrated with Kubernetes (image pull authentication automatic) - Multi-region support - Private registries per project - Pricing: €0.01/GB/month storage **Managed DNS:** - API-driven (works with external-dns, cert-manager) - Support for all record types - DNSSEC available - Pricing: included with domains **Transactional Email (TEM):** - SMTP service for application emails - Scaleway domain verification - Simple integration (SMTP credentials in app configuration) - Pricing: pay per email sent **This combination is what sets Scaleway apart:** It's not just VPS hosting. It's a complete platform where the pieces integrate cleanly. ## How European Providers Compare Based on our 2025 evaluation: **Scaleway:** - Most complete service offering - Strong multi-AZ support (3 AZs per region) - Good Kubernetes integration - Pricing competitive - Support responsive when needed **Best for:** Organizations wanting managed Kubernetes with full ecosystem. **OVHcloud:** - Larger scale, broader European presence - Multiple European regions (France, Germany, Poland, UK) - Managed Kubernetes available - Strong enterprise customer base **Best for:** Larger organizations needing wide European geographic distribution. **Hetzner:** - Excellent price-performance on compute - Less comprehensive managed services - Strong in Germany/Finland - Managed Kubernetes is newer **Best for:** Organizations comfortable managing more themselves, cost-sensitive workloads. **IONOS:** - Strong in Germany - Managed Kubernetes available - Good for German regulatory requirements **Best for:** German-focused workloads with specific compliance needs. ## What We've Learned Running on Scaleway After 18+ months in production: **What works well:** - Kubernetes is solid (control plane reliable, upgrades smooth) - Object storage is reliable (we've had zero data loss incidents) - Cost is predictable and reasonable (no surprise bills) - Support genuinely helps when needed (humans who understand the platform) - API integration works (external-dns, cert-manager, OpenTofu all integrate cleanly) **Where it's maturing:** - Some features feel early-stage (documentation sometimes lags) - Fewer "bells and whistles" than hyperscalers - Some managed services less mature (managed PostgreSQL is good; some others are catching up) - Smaller community (fewer Stack Overflow answers, but growing) **Overall assessment:** It feels like AWS did around 2012—capable, improving rapidly, occasional rough edges. For our workloads, it's a good fit. Importantly: **this assessment could change**. We evaluate annually. If another provider offers better combination of services, we can migrate—because we've built on Kubernetes and portable patterns. ## The "For Now" Mindset We chose Scaleway because it currently meets our requirements. If another provider pulls ahead, we can migrate—because we've built portably. This isn't commitment anxiety; it's good architecture. Provider relationships should be partnerships, not lock-in. **What portability looks like:** - Kubernetes manifests (not provider-specific) - S3-compatible storage (not proprietary APIs) - Standard DNS (RFC-compliant) - SMTP email (standard protocol) - Container registry (standard OCI format) These choices mean: **switching providers is possible**. Not trivial—there's always work involved—but possible. That optionality has value. ## Practical Provider Selection **For Kubernetes-native applications:** Scaleway Kapsule, OVHcloud Managed Kubernetes, IONOS—all offer solid managed Kubernetes. We run production on Scaleway; it's stable, well-priced, and support is responsive. **For compute-intensive work:** Hetzner offers excellent price-performance on dedicated servers. If you can manage your own orchestration, it's hard to beat. **For object storage:** S3-compatible storage is widely available. Scaleway, OVHcloud, Wasabi with European regions—all work with standard S3 tooling. **For specific compliance requirements:** Map your regulatory requirements to provider certifications. SOC 2, ISO 27001, GDPR compliance—verify before committing. ## Making the Choice When evaluating European providers: 1. **Define your requirements**: What do you actually need? (Managed K8s? Multi-AZ? Specific services?) 2. **Map providers to criteria**: Which providers offer what you need? 3. **Test with pilot project**: Don't commit fully. Run one workload for 3 months. Learn what works. 4. **Evaluate honestly**: Provider marketing vs. reality. What actually works in production? 5. **Maintain portability**: Build so you can migrate if needed. This preserves options. The goal isn't finding the perfect provider forever. It's finding the right provider for now, while maintaining ability to change if circumstances change. **Sources:** - [Scaleway Pricing](https://www.scaleway.com/en/pricing/) - [Scaleway Object Storage Documentation](https://www.scaleway.com/en/object-storage/) #scaleway #kubernetes #europeancloud #providerselection #infrastructure --- ## Kubernetes as the Independence Layer URL: https://clouds-of-europe.eu/content/policy-sovereignty/open-source-standards/kubernetes-as-the-independence-layer Author: Jurg van Vliet Published: 2025-07-10 Category: Policy & Sovereignty Type: Open Source & Standards Tags: cloudsofeurope, independence, kubernetes, portability, standards ## What Makes a Standard vs. a Product Before August 2023, many organisations treated Terraform as an open standard. It was widely adopted. It had a permissive license (MPL 2.0). It felt like public infrastructure. Then HashiCorp changed the license to Business Source License (BSL). The community forked it. And we learned the difference between "widely adopted" and "truly open." **Terraform was a product:** One company controlled it. That company could—and did—change the terms. The fork (OpenTofu) was possible because of the previous MPL license, but the disruption was real. **Kubernetes is a standard:** Governed by the CNCF (Cloud Native Computing Foundation), part of the Linux Foundation. No single company can change its terms or direction unilaterally. For foundational infrastructure—the layer everything else depends on—this distinction is critical. ## How Kubernetes Governance Actually Works Kubernetes isn't controlled by Google, despite Google originating the project in 2014. In 2015, Google donated Kubernetes to the CNCF, ceding control to foundation governance. **Decision-making structure:** **Steering Committee:** Representatives from multiple organizations (Google, Red Hat, Microsoft, VMware, independent contributors). Makes architectural decisions through consensus. **Special Interest Groups (SIGs):** Domain-specific groups (SIG Network, SIG Storage, SIG Security) that own specific aspects. Anyone can participate. **Technical Oversight Committee:** Governs technical processes, defines project scope, resolves disputes. **Contributors:** Over 88,000 individuals from more than 8,000 companies across 44 countries. This is the second-largest open source project globally (after Linux kernel). **No single company can unilaterally:** - Change Kubernetes license - Remove features other companies depend on - Force architectural direction - Monetize the project proprietary This distributed governance creates stability. Decisions emerge from consensus, not corporate mandate. ## Multiple Implementations Prove It's a Standard **Managed Kubernetes offerings:** - Google GKE (Google Cloud) - AWS EKS (Amazon) - Azure AKS (Microsoft) - Scaleway Kapsule (Scaleway) - OVHcloud Managed Kubernetes - IONOS Managed Kubernetes - DigitalOcean Kubernetes - Dozens more **Self-hosted distributions:** - OpenShift (Red Hat) - Rancher (SUSE) - K3s (lightweight Kubernetes) - MicroK8s (Canonical) - Talos (immutable infrastructure) All of these implement the same Kubernetes API. A deployment manifest written for GKE works on Scaleway Kapsule. This is portability through standards. ## What True Portability Looks Like Our production deployment: ```yaml apiVersion: apps/v1 kind: Deployment metadata: name: clouds-of-europe-app namespace: app spec: replicas: 3 selector: matchLabels: app: web template: metadata: labels: app: web spec: containers: - name: app image: rg.fr-par.scw.cloud/clouds-of-europe/app:v1.2.3 ports: - containerPort: 3000 env: - name: DATABASE_URL valueFrom: secretKeyRef: name: database-credentials key: url resources: requests: memory: "512Mi" cpu: "250m" limits: memory: "1Gi" cpu: "1000m" livenessProbe: httpGet: path: /api/health port: 3000 initialDelaySeconds: 30 readinessProbe: httpGet: path: /api/health port: 3000 initialDelaySeconds: 5 ``` This is vanilla Kubernetes. It runs on: - Scaleway Kapsule (current production) - OVHcloud Managed Kubernetes (tested) - Kind (local development) - Any Kubernetes 1.28+ cluster **Only provider-specific element:** Image registry URL (`rg.fr-par.scw.cloud`). Change to different registry, manifest still works. We've actually tested this. Same manifest, different clusters, works identically. That's the value of standards. ## What Breaks Portability **Provider-specific annotations:** ```yaml # AWS-specific (breaks portability) metadata: annotations: service.beta.kubernetes.io/aws-load-balancer-type: "nlb" service.beta.kubernetes.io/aws-load-balancer-internal: "true" # GCP-specific (breaks portability) metadata: annotations: cloud.google.com/load-balancer-type: "Internal" cloud.google.com/backend-config: "backend-config" # Scaleway-specific (breaks portability) metadata: annotations: service.beta.kubernetes.io/scw-loadbalancer-proxy-protocol-v2: "true" ``` Every provider-specific annotation is a hidden dependency you'll need to address during migration. **Proprietary storage classes:** ```yaml # AWS-specific EBS storage volumeClaimTemplates: spec: storageClassName: gp3 # AWS EBS gp3 ``` Using cloud-specific storage classes ties you to that provider's storage implementation. Better: use generic storage class names that providers implement consistently. **Custom resource definitions (CRDs):** AWS Controllers for Kubernetes (ACK), Azure Service Operator, GCP Config Connector—these let you manage cloud resources from Kubernetes. They're also completely proprietary. Every resource defined with provider CRDs is something you can't move elsewhere. ## The Portable Architecture Pattern **What we do:** **1. Stick to vanilla Kubernetes APIs** Use Deployment, Service, Ingress (now Gateway API), ConfigMap, Secret—the standard resources. These work everywhere. **2. Abstract provider-specific needs** Need object storage? Use S3-compatible API (works on AWS S3, Scaleway Object Storage, Wasabi, self-hosted MinIO). Need DNS? Use external-dns with standard annotations (works with Route53, Scaleway DNS, Cloudflare). Need certificates? Use cert-manager with ACME or DNS-01 (works with Let's Encrypt on any provider). **3. Use Gateway API for routing** As described in our Envoy migration article—Gateway API provides implementation-agnostic routing. Change proxy implementation without touching HTTPRoute definitions. **4. Document provider-specific choices** When you must use provider-specific features, document it explicitly. Make it a conscious tradeoff, not accidental dependency. ## Why This Matters for Independence **Scenario 1: Locked in** You've built on AWS Lambda, DynamoDB, CloudFormation, SQS. Five years invested. AWS increases pricing 30% at renewal. You can't credibly threaten to leave—migration would cost more than paying the increase. **Scenario 2: Portable** You've built on Kubernetes, PostgreSQL, RabbitMQ, OpenTofu. Five years invested on AWS. AWS increases pricing 30% at renewal. You can credibly say: "We can migrate to Scaleway in 3 months for €X cost. Will you match competitive pricing?" Portability creates **negotiating optionality**. You might never migrate. But the ability to migrate changes pricing conversations. ## The Kubernetes Portability Test Can you rebuild your infrastructure on a different provider in under 1 week with 1 engineer? If yes: you've achieved portability. Congratulations. If no: identify what's preventing it. Provider-specific services? Proprietary APIs? Lack of documentation? Fix these incrementally. Each improvement increases optionality and reduces lock-in risk. **Sources:** - [CNCF: Kubernetes Contributors](https://www.cncf.io/) - [Kubernetes Governance](https://github.com/kubernetes/community/blob/master/governance.md) #kubernetes #portability #standards #independence #cloudsofeurope --- ## Build on Standards, Not Products: Lessons from the Terraform Fork URL: https://clouds-of-europe.eu/content/policy-sovereignty/open-source-standards/build-on-standards-not-products-lessons-from-the-terraform-f Author: Jurg van Vliet Published: 2025-06-12 Category: Policy & Sovereignty Type: Open Source & Standards Tags: kubernetes, opensource, opentofu, portability, standards ## The Terraform Wake-Up Call In August 2023, HashiCorp changed Terraform's license from Mozilla Public License 2.0 (MPL) to Business Source License 1.1 (BSL). For many organisations—including ours—this was clarifying. We'd invested heavily in Terraform. It felt like a standard. But it was a **product controlled by a single company**, and products can change direction. The BSL isn't open source—it's source-available with restrictions. Organisations providing services competitive with HashiCorp could no longer use Terraform freely. The license change was retroactive to all future releases. ## The Community Response August 15, 2023: The OpenTF Manifesto was released, asking HashiCorp to reverse the license change. Over 100 companies, 10 projects, and 400 individuals pledged support. August 25, 2023: OpenTF announced an open source fork of Terraform. September 20, 2023: The Linux Foundation accepted the project as OpenTofu. Within weeks of the license change, the community had forked the project and established it under foundation governance. This speed demonstrates the value the community placed on truly open licensing. ## Our Migration to OpenTofu We migrated to OpenTofu immediately. The transition was straightforward—OpenTofu maintains API compatibility with Terraform. Our configuration files worked unchanged: ```hcl terraform { required_providers { scaleway = { source = "scaleway/scaleway" version = "~> 2.0" } } } ``` Simply changing the binary from `terraform` to `tofu` was sufficient. State files migrated cleanly. Provider plugins worked identically. **The lesson:** There's a difference between "widely adopted" and "truly open." Terraform was the former; Kubernetes is the latter. For foundational infrastructure, this distinction matters. ## What Makes Kubernetes Different Kubernetes is governed by the Cloud Native Computing Foundation, itself part of the Linux Foundation. This provides **structural independence from any single vendor**. **Key differences:** **Governance:** CNCF Technical Oversight Committee makes decisions. Members represent diverse organisations—Google, Red Hat, Microsoft, Huawei, independent contributors. No single company controls direction. **License:** Apache 2.0—a true open source license with no restrictions on use, modification, or commercial deployment. **Multiple implementations:** Kubernetes isn't one product. It's a specification with implementations from Google, Amazon, Microsoft, Red Hat, Rancher, and dozens more. If one implementation becomes problematic, alternatives exist. **Contribution model:** Over 88,000 contributors from 8,000+ companies. The project isn't dependent on any single organization's continued participation. **Cannot be relicensed:** Apache 2.0 contributions cannot be relicensed without contributor consent. The project cannot be "BSL-ed" by a corporate owner. This is what genuine openness looks like: **distributed governance, true open source license, multiple implementations, broad participation**. ## Why Kubernetes Enables Independence Kubernetes provides a consistent API across every major cloud provider and on-premises environment. When you deploy a workload on Kubernetes, you're writing to a standard governed by the CNCF, not a single vendor. This matters practically. We run production on Scaleway Kapsule. Our development environment uses Kind (Kubernetes in Docker). The same manifests work in both places. If we needed to move to OVHcloud, Hetzner Cloud, or even AWS EKS, our application definitions wouldn't change. **What to watch for:** Provider-specific features create hidden dependencies. Every custom annotation, every proprietary storage class, every cloud-specific load balancer integration is something you'll need to address during migration. Example of portable Kubernetes: ```yaml apiVersion: apps/v1 kind: Deployment metadata: name: app spec: replicas: 3 selector: matchLabels: app: web template: spec: containers: - name: app image: myapp:v1 ports: - containerPort: 8080 ``` This manifest works identically on every Kubernetes implementation. That's portability through standards. Example of non-portable Kubernetes: ```yaml apiVersion: v1 kind: Service metadata: name: app annotations: service.beta.kubernetes.io/aws-load-balancer-type: "nlb" service.beta.kubernetes.io/aws-load-balancer-internal: "true" spec: type: LoadBalancer ``` These AWS-specific annotations break portability. When you move to another provider, you'll need to rewrite this configuration. **Stick to vanilla Kubernetes APIs where possible.** Provider-specific features should be explicitly chosen tradeoffs, not accidental dependencies. ## A Checklist for Evaluating Technology Choices Before adopting any technology for your infrastructure, ask: **1. Who controls it?** - Foundation governance (CNCF, Linux Foundation, Apache) ✓ - Multi-company consortium (debatable, depends on structure) - Single company (product, not standard) ✗ **2. What license?** - Apache 2.0, MIT, MPL 2.0 (truly open) ✓ - BSL, SSPL, proprietary (restricted use) ✗ **3. How many implementations?** - Multiple independent implementations (real standard) ✓ - Single implementation (single point of control) ✗ **4. Can it be relicensed?** - Foundation-owned, irrevocable license ✓ - Company-owned, license can change (risk) ✗ **5. What's the exit path?** - Clear migration path to alternatives ✓ - Locked in, migration extremely costly ✗ For foundational infrastructure—the stuff that's expensive to change—prefer technologies that score well on all five criteria. For higher-level tooling, commercial products can be acceptable if you understand the tradeoffs. ## Beyond Kubernetes Other examples of truly open standards: **PostgreSQL:** Open source database with PostgreSQL Global Development Group governance. Implementations include AWS RDS, Azure Database, Google Cloud SQL, self-hosted. License: PostgreSQL License (BSD-like). **Prometheus:** CNCF-graduated monitoring project. Implementations include self-hosted, Grafana Cloud, AWS Managed Prometheus. License: Apache 2.0. **Envoy Proxy:** CNCF-graduated proxy. Implementations include standalone, Istio, AWS App Mesh, Envoy Gateway. License: Apache 2.0. These are genuine standards: **foundation governance, truly open licenses, multiple implementations**. ## The Pattern Build your infrastructure on standards, not products. Products can change direction, relicense, or disappear. Standards—especially those governed by foundations with diverse participation—provide stability. This doesn't mean never use commercial products. It means: **understand what's a standard and what's a product, and choose deliberately**. For the core infrastructure that everything depends on, standards provide the independence that makes everything else possible. **Sources:** - [HashiCorp License Change (August 2023)](https://spacelift.io/blog/terraform-license-change) - [OpenTofu Announces Fork of Terraform](https://opentofu.org/blog/opentofu-announces-fork-of-terraform/) - [The Register: HashiCorp's license shakeup seeded open source rebel](https://www.theregister.com/2024/04/04/opentofu_on_forking_terraform/) - [CNCF: Digital transformation driven by community](https://www.cncf.io/blog/2025/01/30/digital-transformation-driven-by-community-kubernetes-as-example/) #opentofu #kubernetes #standards #opensource #portability --- ## Open Source Mirrors European Governance: Why Standards Fit Our Structure URL: https://clouds-of-europe.eu/content/policy-sovereignty/open-source-standards/open-source-mirrors-european-governance-why-standards-fit-ou Author: Jurg van Vliet Published: 2025-06-08 Category: Policy & Sovereignty Type: Open Source & Standards Tags: cloudsofeurope, europeanvalues, governance, opensource, standards ## Two Models of Governance The European Union represents a governance model that differs fundamentally from centralized nation-states. Multiple sovereign entities collaborate through shared standards and frameworks. No single nation controls the EU; decisions emerge from negotiation and consensus among peers. This isn't always efficient. It's often slower than top-down decision-making. But it creates resilience through diversity: no single point of failure, no single point of control. Open source infrastructure follows remarkably similar patterns. ## How Open Source Governance Works Consider Kubernetes, governed by the Cloud Native Computing Foundation (CNCF): **Multiple implementations**: Kubernetes isn't a single product. It's a specification with implementations from Google (GKE), Amazon (EKS), Microsoft (AKS), Red Hat (OpenShift), Rancher, K3s, and dozens more. No single vendor controls the standard. **Standards bodies**: The CNCF operates through working groups, special interest groups, and technical committees. Decisions emerge from consensus among contributors representing different organizations. Over 88,000 contributors from more than 8,000 companies across 44 countries participate. **Distributed decision-making**: No single company can unilaterally change Kubernetes. Even Google—which originated the project—must work through community governance processes. **Federation over monopoly**: Rather than one massive Kubernetes implementation, the ecosystem consists of many implementations that interoperate through common APIs and standards. ## Contrast with Proprietary Cloud Models US technology platforms tend toward different organizational patterns: **Single platform dominance**: AWS holds roughly 32% of cloud infrastructure market. Within its ecosystem, AWS makes unilateral decisions about features, pricing, deprecations. **Proprietary control**: Vendor lock-in by design. Services like Lambda, DynamoDB, and SQS are available only from AWS. You can't take these workloads elsewhere without rewriting. **Winner-takes-all markets**: Network effects favor the largest platform. The biggest gets more customers, which provides data to build better services, which attracts more customers. **Centralized decision-making**: Corporate hierarchy determines product direction. Customers can request features but can't participate in governance. This model optimizes for efficiency and rapid scaling. It doesn't optimize for customer independence or distributed control. ## Why This Alignment Matters When European organisations build on open source infrastructure, they're working with governance models that mirror their own political structures. This isn't superficial—it's **structural compatibility**. **Both value:** - Distributed sovereignty over centralized control - Standards and interoperability over proprietary lock-in - Consensus decision-making over top-down mandates - Federation over consolidation - Resilience through diversity over efficiency through monopoly European political culture already understands these tradeoffs. We accept that EU decision-making is slower than nation-state unilateral action because we value the resilience that distributed governance provides. The same reasoning applies to technology: we accept that open standards evolve slower than proprietary products because we value the independence that multiple implementations provide. ## Practical Implications **When choosing technology for foundational infrastructure, ask:** 1. **Who governs it?** A foundation (CNCF, Linux Foundation, Apache) or a single company? 2. **How are decisions made?** Community consensus or corporate mandate? 3. **How many implementations exist?** One implementation means one point of control. Multiple implementations mean genuine standard. 4. **Can you participate in governance?** Can your engineers contribute to direction, or just consume what's provided? **For critical infrastructure—the stuff that's expensive to change—prefer truly open governance:** - Kubernetes over proprietary container orchestration - PostgreSQL over proprietary databases - Prometheus over proprietary monitoring - OpenTofu over license-restricted infrastructure tools **For higher-level tooling, commercial products can be fine—if you understand the tradeoffs:** Using proprietary services isn't wrong. But understand what you're trading: convenience and features for increased dependency and reduced optionality. ## The European Opportunity Europe won't out-compete US tech companies by copying their centralized, winner-takes-all model. We don't have the capital, the market scale, or frankly the cultural inclination. But Europe can lead in **federated, standards-based, distributed governance**—because this aligns with our existing political structures and cultural values. Open source infrastructure isn't just compatible with European thinking. It's the natural technical expression of European political philosophy: **multiple sovereign actors collaborating through shared standards**. This is why European organisations building on Kubernetes, PostgreSQL, and open standards aren't just making technical choices. They're building infrastructure that aligns with European values and governance models. **Sources:** - [Cloud Native Computing Foundation](https://www.cncf.io/) - [CNCF: Digital transformation driven by community (2025)](https://www.cncf.io/blog/2025/01/30/digital-transformation-driven-by-community-kubernetes-as-example/) #opensource #europeanvalues #governance #standards #cloudsofeurope --- ## CLOUD Act: Europe pays for its own dependence URL: https://clouds-of-europe.eu/content/policy-sovereignty/policy-regulation/cloud-act-europe-pays-for-its-own-dependence Author: Jurg van Vliet Published: 2025-06-05 Category: Policy & Sovereignty Type: Policy & Regulation A practical question for every European CTO: If a US court issues a subpoena tomorrow for your customer data, would your cloud provider be legally obligated to comply? For most European companies running on AWS, Azure, or Google Cloud, the answer is unequivocally **yes**. The US **CLOUD Act** (Clarifying Lawful Overseas Use of Data Act, 2018\) requires US companies to provide data to US law enforcement, regardless of the data's physical storage location. A server farm in Frankfurt or Stockholm does not change this; what matters is which company controls the keys. This legal reality was solidified by the precedent set during the **Microsoft Ireland Case**. In 2013, Microsoft challenged an FBI warrant demanding emails stored on servers in Dublin. The case reached the US Supreme Court. While arguments were underway, Congress passed the CLOUD Act, explicitly amending the Stored Communications Act to mandate compliance with US law enforcement requests *“regardless of whether such communication, record, or other information is located within or outside of the United States.”* The Supreme Court declared the case moot, establishing a clear framework: **jurisdiction follows the corporate entity, not the server location.** Beyond legal jurisdiction, there is a core economic consideration: **where do the profits, and strategic power, flow?** When a European organisation pays AWS, Azure, or Google Cloud, those revenues ultimately fund US-headquartered corporations. - **AWS Europe (Frankfurt, Paris)**: Owned by Amazon.com, Inc., incorporated in Delaware, USA. - **Azure Europe**: Owned by Microsoft Corporation, incorporated in Washington, USA. - **Google Cloud Europe**: Owned by Google LLC, incorporated in Delaware, USA. These corporations, even operating through European subsidiaries, remit profits back to their US parent companies. This creates a one-way transfer: the vast innovation capacity built from European expenditure flows back to corporate headquarters in Seattle, Redmond, or Mountain View. Investment decisions, core research priorities, and strategic direction for the global cloud market are made there, not in Europe. In contrast, paying European providers like **OVHcloud (French), Scaleway (French), or Hetzner (German)** ensures that revenue remains within the European ecosystem. These companies reinvest in European data centers, support European engineering teams, and contribute to local innovation capacity. This is a pragmatic economic choice: **Are you building digital infrastructure for Europe, or paying a premium to reinforce a non-European strategic lead?** We have reviewed dozens of enterprise cloud agreements. They are sophisticated commercial contracts with detailed Service Level Agreements (SLAs) covering uptime and response times. However, no commercial contract can override the laws governing the provider's home jurisdiction. An SLA can promise 99.99% availability. It cannot promise that your customer data will not be disclosed under a valid US legal order. This creates an acute and uncomfortable situation for European organisations: you are contractually and legally obligated to protect customer data under **GDPR**, while your infrastructure provider is potentially obligated to disclose it under US law. These obligations are in direct conflict. The **Schrems II decision (2020)** already invalidated Privacy Shield and confirmed that US surveillance law poses significant obstacles for EU-US data transfers. Organisations using US cloud providers must implement complex supplementary measures, standard contractual clauses, and transfer impact assessments to manage this risk. For European-controlled infrastructure, these complications are largely eliminated. The provider and your organisation operate under the same legal and regulatory framework. Running on European infrastructure doesn't replace compliance, it simplifies it. Choosing infrastructure that aligns with compliance requirements and strategic independence is good architecture. European cloud providers operate under European jurisdiction, meaning data stored with them is subject to European law. This is the strategic choice for those concerned with **GDPR, NIS2, and the forthcoming AI Act**. **A Pragmatic Action Plan for European Tech Leadership:** 1. **Identify Sensitive Workloads:** Determine which data and applications have the strictest regulatory, contractual, or sovereignty requirements. 2. **Map Current Providers:** Clearly document where each provider is incorporated and which national laws govern them. 3. **Plan Incrementally:** Full migration is not always necessary or feasible. Start by moving the most sensitive and jurisdiction-critical workloads to European-controlled infrastructure. Building on infrastructure you control is not a political statement or an act of nationalism. It is a necessary, fact-based step toward ensuring genuine digital sovereignty and meeting the legal and strategic mandates of a European enterprise. --- ## Sustainability as Europe's Competitive Advantage URL: https://clouds-of-europe.eu/content/policy-sovereignty/resource-efficiency-policy/sustainability-as-europes-competitive-advantage Author: Jurg van Vliet Published: 2025-01-16 Category: Policy & Sovereignty Type: Resource Efficiency (Policy) Tags: efficiency, europe, greencloud, renewable, sustainability While American hyperscalers compete on scale and features, Europe has a unique opportunity to lead in sustainable cloud computing. With ambitious climate goals, renewable energy leadership, and strong environmental regulations, Europe can make sustainability a core competitive advantage in cloud infrastructure. ## The Environmental Imperative Data centers consume about 1% of global electricity—comparable to entire countries like Argentina. As digitalization accelerates, this will only grow. Traditional cloud providers optimize for performance and cost, treating energy efficiency as an afterthought. Europe can flip this model, making sustainability the primary design principle. European advantages: - **Renewable Energy Access**: Nordic countries offer abundant hydroelectric power; Spain and Portugal lead in solar - **Cool Climates**: Natural cooling reduces energy needs in Northern Europe - **Circular Economy Leadership**: European regulations drive hardware recycling and reuse - **Carbon Pricing**: EU emissions trading makes efficiency economically attractive ## Sustainable by Design European cloud infrastructure can pioneer sustainable computing patterns: **Energy-Aware Scheduling**: Workloads run when renewable energy is abundant. AI training happens during sunny afternoons in Spain; batch processing runs during windy nights in Denmark. **Efficient Hardware Utilization**: American clouds optimize for peak performance, leaving servers idle most of the time. European clouds can optimize for efficiency—better to run 100 servers at 80% utilization than 200 at 40%. **Circular Hardware Lifecycle**: Instead of discarding servers after 3-5 years, European providers can pioneer reuse. Older hardware moves from critical workloads to development environments to edge computing. **Heat Recovery**: Data center cooling produces waste heat. Northern European facilities already warm nearby buildings. This can expand—every data center becoming a community heating resource. ## Measuring What Matters Current cloud providers obscure environmental impact. European clouds must provide transparency: - **Real-time Carbon Intensity**: Show CO₂ per compute hour for every region - **Energy Source Disclosure**: Break down renewable vs. fossil fuel usage - **Hardware Lifecycle Tracking**: Report embodied carbon and recycling rates - **Efficiency Metrics**: Publish PUE (Power Usage Effectiveness) and other efficiency measures This transparency enables informed decisions. Organizations can optimize not just for cost and performance but for environmental impact. ## Economic Benefits of Sustainability Sustainability isn't just ethical—it's economically smart: **Lower Operating Costs**: Efficient operations reduce energy bills. Renewable energy provides price stability. **Regulatory Compliance**: As environmental regulations tighten, sustainable infrastructure becomes mandatory. **Talent Attraction**: Developers, especially younger ones, prefer employers with strong environmental commitments. **Customer Preference**: European consumers and businesses increasingly choose sustainable options. ## The Competitive Advantage Sustainability gives European cloud providers unique market positioning: 1. **Differentiation**: While others compete on features, Europe competes on values 2. **Innovation Driver**: Constraints foster creativity and breakthrough solutions 3. **Regulatory Alignment**: Built-in compliance with evolving environmental laws 4. **Economic Efficiency**: Sustainable operations are ultimately cheaper operations 5. **Brand Value**: Association with environmental leadership Europe doesn't need to match American hyperscalers on their terms. By making sustainability the core of cloud infrastructure, Europe can lead the world toward a digital future that's both powerful and responsible. #sustainability #greencloud #europe #renewable #efficiency --- ## Software Distribution: The Missing Piece URL: https://clouds-of-europe.eu/content/practice/expert-insights/software-distribution-the-missing-piece Author: Jurg van Vliet Published: 2025-01-15 Category: Practice Type: Expert Insights Tags: cloudnative, europe, kubernetes, operators, softwaredistribution Fifteen years ago, infrastructure became programmable through APIs. This transformation enabled the cloud revolution, but we've stalled at infrastructure automation. True cloud independence requires solving software distribution—making complex applications as easy to deploy as mobile apps are to install. ## The Current State: Infrastructure as Code Today's "Infrastructure as Code" is mostly templates and configuration. Terraform scripts, Helm charts, and CloudFormation templates are better than manual configuration but fall short of true software distribution. They're like recipes that require a master chef—you need deep expertise to use them effectively. Consider deploying a production database: - Choose from dozens of configuration options - Set up networking and security correctly - Configure backups and test restore procedures - Implement monitoring and alerting - Plan for scaling and upgrades Each organization reinvents these wheels, making similar mistakes and learning similar lessons. This isn't software distribution—it's infrastructure archaeology. ## The Operator Revolution Kubernetes Operators point toward a better future. An Operator encapsulates operational knowledge, turning human expertise into running code. The PostgreSQL Operator doesn't just deploy a database—it embodies years of DBA experience in software form. But current Operators are like smartphone apps before app stores. They exist, they're powerful, but discovering, trusting, and deploying them requires expertise. We need the cloud equivalent of an app store—curated, tested, one-click deployable applications. ## Europe's Distribution Opportunity Software distribution is where Europe can leapfrog current cloud providers. While they're locked into proprietary service models, Europe can build on open standards: **Universal Package Format**: Like mobile apps run on iOS or Android, cloud applications should run on any Kubernetes. **Automated Operations**: Applications should self-manage. Deploy PostgreSQL, and it handles its own backups, scaling, and upgrades. **Declarative Dependencies**: Applications should declare what they need—storage, networking, other services—and the platform provides it. **Lifecycle Management**: Updates, migrations, and eventual decommissioning should be built-in. ## The App Store Model Imagine a European Cloud Application Store: **For Developers**: - Publish applications once, run anywhere in Europe - Built-in billing and metering - Automatic compliance checking - Community ratings and reviews **For Users**: - One-click deployment of complex applications - Guaranteed compatibility with their infrastructure - Transparent pricing and resource usage - Vendor-neutral alternatives to proprietary services **For Europe**: - Sovereignty through software diversity - Innovation ecosystem around open standards - Economic opportunity for European software companies - Reduced dependence on foreign platforms ## The Path Forward Building true software distribution requires coordinated effort: 1. **Standardize Operators**: Create guidelines for well-behaved, portable Operators 2. **Build Marketplaces**: Develop European application marketplaces with curation and testing 3. **Simplify Deployment**: Make deploying complex applications as easy as installing mobile apps 4. **Ensure Portability**: Test applications across multiple European cloud providers 5. **Create Incentives**: Reward developers who build portable, open applications When software distribution is solved, infrastructure becomes commodity. Organizations choose providers based on values—sustainability, locality, support—not lock-in. #softwaredistribution #kubernetes #operators #cloudnative #europe --- ## European Cloud Deployment Models URL: https://clouds-of-europe.eu/content/strategy-transition/strategic-planning/european-cloud-deployment-models Author: Jurg van Vliet Published: 2024-09-25 Category: Strategy & Transition Type: Strategic Planning Tags: cloudstrategy, deployment, europe, hybrid, sovereignty The terms "public," "private," and "hybrid" cloud have been hijacked by marketing departments, obscuring their real meaning and implications for European digital sovereignty. Understanding these deployment models—and their European alternatives—is crucial for organizations planning their cloud independence journey. ## Redefining "Public" Cloud When AWS, Azure, and Google Cloud call themselves "public" clouds, they're using "public" to mean "available to the public," like a public swimming pool. But these are private companies, subject to private interests and foreign governments. For Europe, we need truly public cloud infrastructure—owned by or accountable to the public, operating under European governance and values. True public cloud characteristics: - **Democratic Governance**: Accountable to citizens, not shareholders - **Transparent Operations**: Open about data handling and security practices - **European Jurisdiction**: Subject only to European laws and courts - **Public Interest Focus**: Prioritizing societal benefit over profit maximization Several European initiatives approach this ideal. Gaia-X creates a federated infrastructure with shared governance. National research networks provide computing resources for academia. Municipal data centers serve local government needs. ## Private Cloud: More Than On-Premises Private cloud isn't just servers in your basement. It's about control—over data, operations, and destiny. Modern private clouds built on Kubernetes and OpenStack provide cloud-like experiences while maintaining complete sovereignty. European organizations are pioneering innovative private cloud models: **Industry Collaboratives**: German automotive companies share private cloud infrastructure for non-competitive functions. This spreads costs while maintaining control. **Regional Clouds**: Nordic countries collaborate on shared infrastructure serving government and healthcare. Data never leaves the region; governance remains local. **Sovereign Stacks**: French organizations use "Cloud de Confiance" certified providers—technically private clouds operated by trusted European companies. ## Hybrid Reality Pure public or private clouds are rare. Most organizations operate hybrid environments, and this isn't a transitional state—it's the end goal. Hybrid cloud lets organizations optimize for different requirements: - **Sovereignty for Sensitive Data**: Personal data, trade secrets, and government information stay on European infrastructure - **Scale for Public Services**: Web frontends and mobile apps leverage global CDNs and edge networks - **Specialization for Specific Needs**: AI training on specialized hardware, archival storage on cost-optimized systems - **Resilience Through Diversity**: Multiple providers prevent single points of failure ## European Deployment Innovations Europe is pioneering new deployment models that transcend traditional categories: **Federated Clouds**: Multiple providers collaborate while maintaining independence. Customers get unified experience; providers keep sovereignty. **Edge-to-Cloud Continuum**: Processing happens wherever it makes sense—sensors, edge devices, regional data centers, central clouds. **Regulatory Clouds**: Infrastructure designed for specific regulations. A GDPR-compliant cloud handles personal data; a financial cloud meets banking requirements. **Green Clouds**: Data centers powered entirely by renewable energy, with workloads scheduled based on energy availability. ## The Portable Future The ultimate goal isn't choosing between public, private, or hybrid—it's making the choice irrelevant. With cloud-native applications on Kubernetes, organizations can move workloads based on changing requirements. This portability is Europe's strategic advantage. While others lock customers into proprietary platforms, Europe can build on open standards that preserve choice. #cloudstrategy #deployment #hybrid #sovereignty #europe --- ## Building New vs. Transforming Existing Systems URL: https://clouds-of-europe.eu/content/practice/implementation-patterns/building-new-vs-transforming-existing-systems Author: Jurg van Vliet Published: 2024-06-27 Category: Practice Type: Implementation Patterns Tags: architecture, cloudnative, kubernetes, migration, transformation Every organization faces a fundamental choice: build new cloud-native systems from scratch (greenfield) or transform existing systems (brownfield). This choice shapes everything from architecture decisions to team structure. Understanding both approaches is crucial for Europe's cloud independence journey. ## The Greenfield Dream Starting fresh with a greenfield project is every engineer's dream. No legacy code, no technical debt, no compromises. You can choose the best technologies and build exactly what you need. For a new IoT platform serving agriculture, for example, you might choose: - **Kubernetes** for orchestration, providing portability and scalability - **Apache Kafka** for event streaming, replacing proprietary message queues - **TimescaleDB** for time-series data, avoiding vendor-specific databases - **Prometheus and Grafana** for monitoring, using open-source standards - **K3s on edge devices** for distributed computing at sensor locations This architecture would be cloud-agnostic, running equally well on OVHcloud, Scaleway, or your own infrastructure. ## The Brownfield Reality Consider a real-world IoT platform with 400 customers, processing billions of events monthly. It runs on AWS using EC2, DynamoDB, SQS, and various other proprietary services. The team knows these services intimately. The platform is reliable and profitable. But it's also completely locked into AWS. Transforming such a system requires careful planning: **Phase 1: Containerize and Standardize** Start with stateless services. Move Python and Scala workers from EC2 to containers. Deploy them on EKS (Amazon's Kubernetes) initially. **Phase 2: Replace Commoditized Services** Tackle standardized services next. Move from Amazon Elasticsearch to the Elasticsearch Operator on Kubernetes. Replace CloudWatch with Prometheus and Grafana. **Phase 3: Strategic Service Migration** The hard part: replacing services like DynamoDB and SQS. Move from SQS to Apache Kafka—not just a queue replacement but an opportunity for new event-driven features. **Phase 4: Infrastructure Portability** Once services are containerized and using open standards, you can move workloads. Start with development environments on European clouds. ## Patterns for Success Whether greenfield or brownfield, successful transformations share patterns: **Start Small, Think Big**: Begin with pilot projects that prove concepts. Design for the end state but implement incrementally. **Invest in Knowledge**: Cloud-native transformation is as much about people as technology. Train teams, hire expertise, engage partners. **Maintain Business Continuity**: Never risk the business for technical purity. Keep systems running while transforming them. ## The Hybrid Future The greenfield/brownfield distinction is becoming less relevant. Modern architectures support gradual transformation. You can run cloud-native workloads alongside legacy systems, moving functionality piece by piece. For Europe, this means we don't need to abandon existing investments to achieve cloud independence. We can transform gradually, maintaining business continuity while building sovereignty. #cloudnative #transformation #kubernetes #migration #architecture --- ## The Uneven Distribution of Cloud Innovation URL: https://clouds-of-europe.eu/content/practice/expert-insights/the-uneven-distribution-of-cloud-innovation Author: Jurg van Vliet Published: 2024-06-05 Category: Practice Type: Expert Insights Tags: adoption, cloudnative, europe, innovation, kubernetes William Gibson famously said, "The future is already here—it's just not evenly distributed." This perfectly describes the current state of cloud-native technology in Europe. While some organizations run cutting-edge Kubernetes platforms rivaling anything in Silicon Valley, others struggle with basic cloud adoption. ## Pockets of Excellence Europe has remarkable examples of cloud-native excellence. CERN, the European Organization for Nuclear Research, runs one of the world's most sophisticated Kubernetes deployments. They process petabytes of data from the Large Hadron Collider using cloud-native technologies. Financial institutions in London, Frankfurt, and Amsterdam run trading platforms on Kubernetes that handle billions in transactions daily. These systems demand microsecond latency, absolute reliability, and strict regulatory compliance—proving that Kubernetes can meet the most demanding requirements. German automotive companies use Kubernetes for everything from factory automation to connected car services. They're pioneering edge computing patterns that will define the next generation of industrial IoT. ## The Adoption Gap Yet for every CERN or cutting-edge bank, there are hundreds of organizations still running traditional infrastructure. Small and medium enterprises (SMEs) that form the backbone of European economies often lack the resources and expertise for cloud-native transformation. This gap isn't just about technology—it's about knowledge, culture, and ecosystem support. Organizations that successfully adopt cloud-native technologies typically have: - Leadership that understands and champions transformation - Access to skilled engineers or capable partners - A culture that embraces experimentation and accepts failure - Clear business drivers for change ## Accelerating Distribution To achieve cloud independence, Europe must accelerate the distribution of cloud-native innovation: 1. **Lower Barriers**: Make cloud-native technologies more accessible through managed services and simplified tools 2. **Share Success Stories**: Publicize European organizations succeeding with cloud-native approaches 3. **Build Communities**: Foster local meetups and online communities where practitioners share knowledge 4. **Incentivize Adoption**: Use public funding and procurement to encourage cloud-native transformation 5. **Create Bridges**: Connect advanced practitioners with organizations beginning their journey The future of European cloud independence is already here in pockets of excellence across the continent. Our challenge is to distribute this future more evenly. #cloudnative #innovation #europe #kubernetes #adoption --- ## Kubernetes: Europe's Path to Cloud Independence URL: https://clouds-of-europe.eu/content/policy-sovereignty/open-source-standards/kubernetes-europes-path-to-cloud-independence Author: Jurg van Vliet Published: 2024-05-28 Category: Policy & Sovereignty Type: Open Source & Standards Tags: cloudnative, europe, independence, kubernetes, opensource In the 1990s, Java promised "Write Once, Run Anywhere"—a vision of software that could run on any platform without modification. That dream failed for desktop applications but is being realized today through Kubernetes. This open-source container orchestration platform represents Europe's best path to cloud independence. ## The Kubernetes Revolution Kubernetes abstracts away the underlying infrastructure, providing a consistent platform whether you're running on AWS, OVHcloud, or your own servers. This abstraction is revolutionary because it breaks vendor lock-in at the infrastructure level. Applications designed for Kubernetes can move between providers with minimal modification. More importantly, Kubernetes has become the de facto standard for cloud-native applications. Every major cloud provider offers a Kubernetes service. This universality means that investing in Kubernetes skills and applications is investing in portable, future-proof technology. ## Beyond Basic Infrastructure Early cloud adoption focused on basic services: virtual machines, storage, databases. But Kubernetes enables something more profound through the Operator pattern. Operators are applications that extend Kubernetes to manage complex stateful services like databases, message queues, and monitoring systems. This means we can package not just applications but entire operational knowledge into distributable software. A PostgreSQL Operator doesn't just deploy a database—it handles backups, failovers, scaling, and upgrades. ## Europe's Kubernetes Advantage Europe has several advantages in the Kubernetes ecosystem: **Strong Open Source Culture**: European developers and companies are major contributors to Kubernetes and related projects. This gives Europe influence over the platform's direction and deep expertise in its use. **Regulatory Alignment**: Kubernetes' open governance model aligns with European values of transparency and democratic decision-making. No single company controls it. **Existing Providers**: European cloud providers like OVHcloud, Scaleway, and IONOS already offer managed Kubernetes services. The foundation exists—it needs expansion and integration. ## The Path Forward Kubernetes provides the technical foundation for European cloud independence. But technology alone isn't enough. We need: 1. **More Operators**: European companies should contribute Operators for their specialized needs 2. **Better Integration**: Kubernetes services should integrate seamlessly across European providers 3. **Simplified Experiences**: Developer-friendly tools that hide Kubernetes complexity 4. **Shared Standards**: Common approaches to security, compliance, and operations Kubernetes isn't just another technology choice—it's the key to breaking free from vendor lock-in while maintaining the innovation and convenience that made cloud computing transformative. #kubernetes #cloudnative #independence #opensource #europe --- ## European Organizations and Their Cloud Journey URL: https://clouds-of-europe.eu/content/strategy-transition/success-stories/european-organizations-and-their-cloud-journey Author: Jurg van Vliet Published: 2024-05-22 Category: Strategy & Transition Type: Success Stories Tags: cloudadoption, digitaltransformation, europe, sovereignty, startups To understand where European cloud infrastructure needs to go, we must first understand where European organizations are today. The cloud adoption stories of European companies reveal both the transformative power of cloud computing and the urgent need for sovereign alternatives. ## The Startup Perspective Consider companies like 30MHz, which built IoT platforms for precision agriculture, or Quatt, developing smart heat pumps for the energy transition. These startups couldn't exist without cloud infrastructure. With teams of 3-4 developers, they serve hundreds of customers and manage thousands of devices. These startups chose American cloud providers not out of preference but out of necessity. They needed: - Instant scalability to handle growth - Managed services to focus on their core product - Global presence to serve international customers - Pay-as-you-go pricing to manage cash flow ## Enterprise Transformation Large European enterprises tell a different story. Educational publisher Malmberg moved to AWS to achieve stability and focus on serving students rather than managing servers. When acquired by Sanoma Learning, this cloud-first approach spread across the organization. Financial services platform Ohpen pioneered cloud adoption in the heavily regulated banking sector, proving that even the most compliance-heavy industries could benefit from cloud transformation. ## Public Sector Challenges Perhaps most telling is the story of SIDN, the organization managing the .nl domain. Even this critical piece of Dutch internet infrastructure is moving to AWS because the specialized services they need simply don't exist elsewhere. ## Lessons for European Cloud These stories teach us what European cloud infrastructure must provide: 1. **Developer-First Design**: If developers don't want to use it, organizations won't adopt it 2. **Comprehensive Services**: Basic infrastructure isn't enough—we need the full stack 3. **Migration Paths**: Organizations need clear paths from their current providers 4. **Reliability at Scale**: European providers must match hyperscaler reliability 5. **Competitive Pricing**: Sovereignty can't come at prohibitive cost #cloudadoption #europe #digitaltransformation #sovereignty #startups --- ## Understanding Cloud Infrastructure in Europe URL: https://clouds-of-europe.eu/content/practice/expert-insights/understanding-cloud-infrastructure-in-europe Author: Jurg van Vliet Published: 2024-04-30 Category: Practice Type: Expert Insights Tags: cloudinfrastructure, digitalsovereignty, europe, independence, kubernetes The term "public cloud" has become synonymous with American hyperscalers, but this naming convention obscures a critical issue: these aren't truly public infrastructures but private platforms controlled by foreign corporations. For Europe to achieve digital sovereignty, we must first understand what cloud infrastructure really is and why it matters. ## The Infrastructure Stack Infrastructure as a Service (IaaS) forms the foundation of modern digital services. It provides the basic building blocks—compute power, storage, networking—that enable everything from simple websites to complex AI applications. When we talk about "taking back the cloud," we're talking about ensuring these fundamental building blocks are available under European control. The evolution from IaaS to Platform as a Service (PaaS) and Software as a Service (SaaS) shows how infrastructure providers gradually moved up the stack, offering increasingly sophisticated services. What started as virtual machines and storage buckets evolved into managed databases, AI services, and complete development platforms. This progression created powerful lock-in effects—the more specialized services you use, the harder it becomes to leave. ## Why Infrastructure Sovereignty Matters Control over digital infrastructure isn't just about data location—it's about economic independence, regulatory compliance, and strategic autonomy. When European companies rely entirely on non-European infrastructure: - They're subject to foreign laws and regulations (like the US CLOUD Act) - Their data can be accessed by foreign governments - They're vulnerable to geopolitical tensions and trade disputes - Innovation and value creation benefits flow outside Europe ## Europe's Infrastructure Landscape Europe isn't starting from zero. Providers like OVHcloud, Scaleway, and Hetzner offer competitive infrastructure services. National initiatives in Germany (Gaia-X), France (Cloud de Confiance), and other countries are building sovereign cloud capabilities. What Europe needs isn't a copy of AWS or Azure, but something better: a cloud infrastructure that combines the innovation and ease-of-use of American platforms with European values of privacy, sustainability, and democratic governance. #cloudinfrastructure #digitalsovereignty #europe #kubernetes #independence --- ## The Road Ahead for European Cloud Independence URL: https://clouds-of-europe.eu/content/strategy-transition/leadership-guides/the-road-ahead-for-european-cloud-independence Author: Jurg van Vliet Published: 2024-04-23 Category: Strategy & Transition Type: Leadership Guides Tags: cloudindependence, europe, future, kubernetes, sovereignty After exploring the challenges and opportunities of European cloud independence, the path forward is clear but demanding. Success requires coordinated action across technology development, policy making, business investment, and cultural change. This isn't a project for one company or country—it's a continental imperative. ## The Four Pillars of Action **1. Distribute More Software** The Kubernetes Operator ecosystem needs explosive growth. Every successful deployment should become a reusable Operator. European organizations must contribute their solutions back to the community. We need: - Financial services Operators meeting European banking regulations - Healthcare Operators compliant with patient privacy laws - Government Operators with built-in audit trails - Industrial Operators for manufacturing and logistics When every organization can deploy sophisticated applications without reinventing wheels, we achieve true software distribution. **2. Create More Kubernetes Clouds** European cloud providers must expand beyond basic infrastructure. Every European country needs at least one sovereign Kubernetes provider. Regional providers should federate for resilience. Key requirements: - Managed Kubernetes matching EKS/GKE capabilities - Integration with European identity providers - Built-in compliance for GDPR and sector regulations - Sustainable operations with transparent metrics **3. Build the Management Layer** Hybrid and multi-cloud reality requires sophisticated management. European organizations need tools that work across providers—European and global. This management layer must provide: - Unified visibility across all infrastructure - Consistent security and compliance policies - Cost optimization across providers - Application portability without vendor tools **4. Expose Sustainability Metrics** What gets measured gets managed. Every cloud operation should expose: - Real-time energy consumption and carbon intensity - Hardware utilization and efficiency metrics - Renewable energy percentage - Circular economy indicators ## The European Advantage Europe has unique strengths in this journey: **Regulatory Leadership**: GDPR showed the world how to protect privacy. Similar leadership in cloud sovereignty can set global standards. **Open Source Culture**: European developers contribute significantly to open-source projects. This collaborative culture is perfect for building shared infrastructure. **Sustainability Commitment**: European climate goals create natural alignment between cloud efficiency and policy objectives. **Diverse Economy**: From German manufacturing to French luxury to Nordic design, Europe's economic diversity demands flexible infrastructure. ## The 2030 Vision By 2030, Europe can achieve true cloud independence: - **Technical Sovereignty**: Critical infrastructure runs on European-controlled platforms - **Economic Vibrancy**: Thousands of European cloud companies serve millions of customers - **Environmental Leadership**: European clouds set global standards for sustainable computing - **Democratic Values**: Cloud infrastructure reflects European principles of privacy and governance - **Global Competitiveness**: European organizations innovate freely without foreign dependencies ## The Call to Action Cloud independence won't happen automatically. It requires conscious choice and sustained effort. Every European organization faces a decision: continue down the path of foreign dependency or invest in sovereign alternatives. The technology exists. Kubernetes provides the platform. European providers offer infrastructure. What's needed now is collective will—developers choosing open standards, organizations demanding portability, providers collaborating on ecosystems, and policymakers supporting sovereignty. The clouds of Europe are gathering. It's time to make them rain innovation. #cloudindependence #europe #kubernetes #sovereignty #future ---