GKE — Google Kubernetes Engine
GKE is GCP's managed Kubernetes. Google invented Kubernetes, so GKE gets features before any other cloud (Autopilot, Workload Identity, GKE Gateway, etc.).
GKE Modes
graph TD
classDef cp fill:#4285F4,stroke:#1a73e8,color:#fff
classDef standard fill:#34A853,stroke:#188038,color:#fff
classDef autopilot fill:#8e44ad,stroke:#6c3483,color:#fff
classDef limit fill:#EA4335,stroke:#c5221f,color:#fff
subgraph STANDARD["GKE Standard — you manage nodes"]
STD_CP["Control plane<br/>Google managed"]:::cp
STD_NG["Node groups — you provision<br/>choose machine type, disk, OS<br/>you patch the node OS<br/>you set autoscaling bounds"]:::standard
STD_DS["DaemonSets: fully supported<br/>Host networking: yes<br/>SSH to nodes: yes"]:::standard
STD_CP --> STD_NG --> STD_DS
end
subgraph AUTOPILOT["GKE Autopilot — Google manages everything"]
AUTO_CP["Control plane<br/>Google managed"]:::cp
AUTO_N["Nodes — Google managed<br/>auto-provisioned per pod request<br/>you never see or SSH to a node<br/>billed per pod CPU/memory, not per node"]:::autopilot
AUTO_LIMIT["No DaemonSets<br/>No privileged containers<br/>No host networking<br/>Pod Security Standards enforced"]:::limit
AUTO_CP --> AUTO_N --> AUTO_LIMIT
end
| Feature | Standard | Autopilot |
|---|---|---|
| Node management | You | |
| Billing | Per node (whether idle or not) | Per pod (only what you use) |
| DaemonSets | Yes | No |
| Privileged pods | Yes | No |
| Custom node images | Yes | No |
| Best for | Complex workloads, AI/GPU, custom OS | Typical microservices, cost efficiency |
On GKE Autopilot, a node backing your workloads is only 40% utilized by pod requests. Are you billed for the other 60%, the way you would be on Standard?
Control Plane
graph TD
classDef mgmt fill:#4285F4,stroke:#1a73e8,color:#fff
classDef customer fill:#34A853,stroke:#188038,color:#fff
subgraph GMANAGED["Google-managed control plane (Google's VPC)"]
CP["GKE Control Plane<br/>$0.10/cluster/hr (~$73/mo)<br/>Standard & Autopilot alike<br/>1 free zonal cluster per billing account"]:::mgmt
APISERVER["kube-apiserver<br/>kubectl connects via<br/>private or public endpoint"]:::mgmt
ETCD["etcd<br/>Google manages backups<br/>HA replicated across zones (regional)"]:::mgmt
SCHED["kube-scheduler"]:::mgmt
CM["kube-controller-manager<br/>+ cloud-controller-manager<br/>provisions Load Balancers, PVs"]:::mgmt
CP --> APISERVER
CP --> ETCD
CP --> SCHED
CP --> CM
end
subgraph CUSTVPC["Your VPC"]
NG["Node groups<br/>Compute Engine VMs<br/>running kubelet + containerd"]:::customer
end
APISERVER -->|"node pool<br/>registration & control"| NG
kubectl apply, no scaling, no rescheduling — but Pods and Services that are already running keep serving traffic, since the data plane doesn't depend on a live control plane to keep forwarding packets.
Cost, either way: the $0.10/cluster/hr (~$73/mo) management fee applies to all clusters — Standard or Autopilot, zonal or regional — beyond the one free zonal cluster per billing account. Regional no longer carries a separate surcharge; you only pay more because it runs 3 control-plane replicas' worth of Google-managed infrastructure (included in the flat fee) plus your own multi-zone node/compute costs.
A zonal cluster's control-plane zone goes down. Do the Pods that were already running on your nodes go down with it?
Workload Identity — The Right Way to Access GCP APIs
Never put service account keys in pods. Workload Identity maps a K8s ServiceAccount to a GCP IAM Service Account.
sequenceDiagram
autonumber
participant POD as Pod (K8s SA: my-app)
participant MDS as Metadata Server (169.254.169.254)
participant STS as GCP STS
participant API as Cloud Storage / BigQuery
rect rgb(30, 60, 110)
Note over POD,MDS: Trust boundary 1 — inside the cluster, no long-lived key ever leaves here
POD->>MDS: GET /computeMetadata/v1/instance/service-accounts/default/token
MDS->>MDS: mint a short-lived, cluster-signed<br/>Kubernetes projected token for "my-app"
end
rect rgb(90, 60, 20)
Note over MDS,STS: Trust boundary 2 — GCP verifies the binding, not a static credential
MDS->>STS: exchange K8s projected token for a GCP access token
Note over MDS,STS: K8s SA "my-app" is bound to GCP SA "my-app@project.iam"<br/>via roles/iam.workloadIdentityUser
STS-->>MDS: GCP access token, 1 hour TTL
end
MDS-->>POD: access token
POD->>API: API call with access token
Note right of API: API enforces whatever IAM roles<br/>are granted to my-app@project.iam, nothing more
roles/iam.workloadIdentityUser to the specific PROJECT.svc.id.goog[NAMESPACE/KSA_NAME] identity — no key file is ever created or shipped to the pod.
iam.gke.io/gcp-service-account=my-app-sa@PROJECT.iam.gserviceaccount.com tells GKE which GCP identity this KSA is allowed to impersonate.
Why is Workload Identity considered strictly safer than mounting a downloaded GCP service-account key as a Kubernetes Secret?
# Setup Workload Identity
gcloud iam service-accounts create my-app-sa
# Bind K8s SA → GCP SA
gcloud iam service-accounts add-iam-policy-binding \
my-app-sa@PROJECT.iam.gserviceaccount.com \
--role roles/iam.workloadIdentityUser \
--member "serviceAccount:PROJECT.svc.id.goog[NAMESPACE/KSA_NAME]"
# Annotate the K8s ServiceAccount
kubectl annotate serviceaccount KSA_NAME \
iam.gke.io/gcp-service-account=my-app-sa@PROJECT.iam.gserviceaccount.com
GKE Networking — VPC-Native (Alias IPs)
graph TD
classDef vpc fill:#4285F4,stroke:#1a73e8,color:#fff
classDef node fill:#34A853,stroke:#188038,color:#fff
classDef pod fill:#FBBC04,stroke:#f9ab00,color:#000
subgraph VPCNET["VPC: 10.0.0.0/8 — pod IPs are real, routable VPC addresses"]
subgraph CLUSTER["GKE cluster"]
NODE1["Node: 10.128.0.2<br/>Pod CIDR (alias IP range): 10.4.0.0/24"]:::node
NODE2["Node: 10.128.0.3<br/>Pod CIDR (alias IP range): 10.4.1.0/24"]:::node
POD1["Pod: 10.4.0.5"]:::pod
POD2["Pod: 10.4.0.6"]:::pod
POD3["Pod: 10.4.1.5"]:::pod
POD1 --> NODE1
POD2 --> NODE1
POD3 --> NODE2
end
end
OTHERVM["Any other VM in the VPC<br/>no overlay, no VXLAN decapsulation needed"]:::vpc
OTHERVM -.->|"routes directly to<br/>10.4.0.5, no NAT"| POD1
Alias IPs: Pod IPs are real VPC IPs — no overlay network, no encapsulation. Pods are directly routable from any VM in the VPC. No VXLAN overhead.
This is different from EKS: EKS VPC CNI also gives pods real VPC IPs (similar), but the default pod CIDR is a secondary range rather than alias IPs.
A plain Compute Engine VM elsewhere in the same VPC wants to send a packet straight to a Pod IP. Does it need to go through a VXLAN tunnel or a NAT hop first?
Private Clusters
In a private cluster, nodes get only internal IPs (no external IP). The control plane is reachable via a private endpoint — an RFC 1918 address inside your VPC. The control plane itself still runs in a Google-managed VPC that is VPC-peered to yours; the peering is what carries traffic between your nodes and the private control-plane endpoint.
graph TD
subgraph GVPC["Google-managed VPC (control plane)"]
CP["kube-apiserver<br/>private endpoint: 10.0.0.2<br/>inside --master-ipv4-cidr /28<br/>optional public endpoint"]
end
subgraph YOURVPC["Your VPC"]
subgraph SUBNET["Node subnet (private)"]
N1["Node: 10.128.0.2<br/>internal IP only"]
N2["Node: 10.128.0.3<br/>internal IP only"]
end
NAT["Cloud NAT<br/>egress for nodes: pull images, reach APIs"]
end
subgraph EXTERNAL["Outside both VPCs"]
INTERNET["Internet<br/>registry.k8s.io, gcr.io, etc."]
ADMIN["Admin / CI<br/>authorized network CIDR"]
end
CP -->|"VPC Peering"| SUBNET
N1 --> NAT
N2 --> NAT
NAT -->|"outbound only"| INTERNET
ADMIN -.->|"only if public endpoint enabled"| CP
classDef google fill:#4285F4,stroke:#1a73e8,color:#fff
classDef yours fill:#34A853,stroke:#188038,color:#fff
classDef ext fill:#EA4335,stroke:#c5221f,color:#fff
class CP,GVPC google
class N1,N2,NAT,SUBNET yours
class INTERNET,ADMIN,EXTERNAL ext
Key flags:
--enable-private-nodes— nodes get internal IPs only (no external IP). Required for a private cluster.--enable-private-endpoint— disables the public control-plane endpoint entirely;kubectlmust originate from inside the VPC (or a connected network via VPN/Interconnect/peering). Omit it to keep a public endpoint locked down by authorized networks.--master-ipv4-cidr— a dedicated /28 for the control plane's private endpoint in the Google-managed VPC. Must not overlap any of your VPC ranges. (VPC-native/alias-IP clusters no longer strictly require this for private endpoints on newer versions, but it is still the canonical way to pin the range.)--master-authorized-networks— allowlist of CIDRs permitted to reach the control-plane endpoint. Essential when a public endpoint is left enabled; restricts who can hit the API server.
Cloud NAT is mandatory for egress: since nodes have no external IP, they cannot pull images from public registries or reach external APIs without a NAT. Attach a Cloud NAT to the node subnet's region. Traffic to Google APIs (gcr.io/Artifact Registry, logging, etc.) can alternatively use Private Google Access on the subnet, avoiding NAT for Google endpoints.
What's the actual difference in effect between --enable-private-nodes and --enable-private-endpoint?
# Private cluster: private nodes + public endpoint locked to authorized networks
gcloud container clusters create private-cluster \
--region us-central1 \
--enable-ip-alias \
--enable-private-nodes \
--master-ipv4-cidr 172.16.0.0/28 \
--master-authorized-networks 203.0.113.0/24,10.0.0.0/8 \
--enable-master-authorized-networks
# Fully private (no public endpoint at all — kubectl only from inside the VPC)
gcloud container clusters create fully-private \
--region us-central1 \
--enable-ip-alias \
--enable-private-nodes \
--enable-private-endpoint \
--master-ipv4-cidr 172.16.0.0/28
# Cloud NAT so private nodes can reach the internet (image pulls, etc.)
gcloud compute routers create nat-router --region us-central1 --network default
gcloud compute routers nats create gke-nat \
--router nat-router --region us-central1 \
--nat-all-subnet-ip-ranges --auto-allocate-nat-external-ips
vs EKS private endpoint: EKS exposes the same public/private toggle via endpointPublicAccess / endpointPrivateAccess, and public access is scoped with publicAccessCidrs (the EKS analog of --master-authorized-networks). The big structural difference: EKS reaches the private API server through cross-account ENIs injected into your subnets, whereas GKE uses VPC Peering to a Google-managed control-plane VPC and needs the dedicated /28 (--master-ipv4-cidr). On both, private-only nodes need managed egress — Cloud NAT on GKE, a NAT Gateway on EKS.
GKE Load Balancing
graph LR
classDef entry fill:#4285F4,stroke:#1a73e8,color:#fff
classDef lb fill:#FBBC04,stroke:#f9ab00,color:#000
classDef neg fill:#8e44ad,stroke:#6c3483,color:#fff
classDef pod fill:#34A853,stroke:#188038,color:#fff
DNS2["DNS: app.example.com"]:::entry --> GLB
subgraph GCP["GCP Load Balancers"]
GLB["Global External LB (Ingress / Gateway)<br/>Anycast IP, HTTP/HTTPS<br/>Cloud Armor WAF, Cloud CDN"]:::lb
RLNLB["Regional Internal LB<br/>ClusterIP or Internal LB<br/>within VPC only"]:::lb
PASSTHROUGH["External TCP/UDP LB<br/>LoadBalancer Service<br/>L4 pass-through"]:::lb
end
subgraph BACKENDS["Backends — container-native (NEG)"]
NEG["Network Endpoint Group<br/>direct pod IPs — no kube-proxy, no NodePort hop"]:::neg
POD4["Pod 10.4.0.5:8080"]:::pod
POD5["Pod 10.4.1.5:8080"]:::pod
NEG --> POD4
NEG --> POD5
end
GLB --> NEG
Container-native load balancing (NEG): GCP LB talks directly to pod IPs — bypasses kube-proxy and NodePort entirely. Lower latency, better health checking, pods visible directly in the GCP console.
cloud.google.com/neg: '{"ingress": true}', telling GKE's NEG controller to manage endpoints for this backend itself instead of relying on NodePort.
pod-ip:port directly to the NEG. No NodePort, no kube-proxy iptables/IPVS rule is involved at any point.
With container-native load balancing (NEG) enabled, does a request from the GCP Load Balancer pass through kube-proxy on its way to a Pod?
# Ingress with container-native LB
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
annotations:
kubernetes.io/ingress.class: "gce"
cloud.google.com/neg: '{"ingress": true}' # enable NEG
spec:
rules:
- host: app.example.com
http:
paths:
- path: /*
backend:
service:
name: my-service
port:
number: 8080
Gateway API
The GKE Gateway API is the successor to Ingress — a Kubernetes-native, role-oriented API for L4/L7 routing. GKE ships built-in GatewayClasses that each map to a specific Google Cloud Load Balancer, so choosing a class is really choosing an LB flavor.
graph TD
classDef platform fill:#4285F4,stroke:#1a73e8,color:#fff
classDef appteam fill:#34A853,stroke:#188038,color:#fff
classDef lb fill:#FBBC04,stroke:#f9ab00,color:#000
subgraph PLATFORM["Owned by platform / infra team"]
GC["GatewayClass<br/>cluster-scoped, infra provided"]:::platform
GW["Gateway<br/>listeners, ports, TLS, LB"]:::platform
GC --> GW
end
subgraph APPTEAM["Owned by app team — can live in app namespaces"]
HR["HTTPRoute<br/>hostnames, paths, header rules, weights"]:::appteam
SVC["Service / NEG<br/>pod IPs"]:::appteam
HR --> SVC
end
GW --> HR
subgraph FLAVORS["GatewayClass → Google Cloud LB flavor"]
GLOBAL["gke-l7-global-external-managed<br/>Global External ALB, Anycast"]:::lb
REGIONAL["gke-l7-regional-external-managed<br/>Regional External ALB"]:::lb
RILB["gke-l7-rilb<br/>Regional Internal ALB, VPC-internal"]:::lb
end
GC -.-> GLOBAL
GC -.-> REGIONAL
GC -.-> RILB
| GatewayClass | Google Cloud LB | Scope |
|---|---|---|
gke-l7-global-external-managed |
Global External Application LB (Anycast IP) | Global, internet-facing |
gke-l7-regional-external-managed |
Regional External Application LB | Single region, internet-facing |
gke-l7-rilb |
Regional Internal Application LB | VPC-internal only |
Advantages over Ingress:
- Role separation:
Gateway(owned by platform/infra — listeners, certs, LB) is decoupled fromHTTPRoute(owned by app teams — routes). Ingress crammed both into one object plus vendor annotations. - Richer routing: native header/query matching, path rewrites, request/response header mutation, and traffic weighting (canary/blue-green) — no annotation soup.
- Cross-namespace & multi-route: many
HTTPRoutes can attach to oneGateway; routes can live in app namespaces.
Under the Gateway API's role split, when an app team wants to shift canary traffic weight from 90/10 to 50/50, do they need to touch the Gateway object owned by the platform team?
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
name: external-gw
spec:
gatewayClassName: gke-l7-global-external-managed
listeners:
- name: https
protocol: HTTPS
port: 443
tls:
mode: Terminate
certificateRefs:
- name: app-tls-cert # Secret or Google-managed cert
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: app-route
spec:
parentRefs:
- name: external-gw
hostnames:
- "app.example.com"
rules:
- matches:
- path:
type: PathPrefix
value: /v2
headers:
- name: x-canary
value: "true"
backendRefs:
- name: app-v2
port: 8080
- matches: # weighted split for the rest
- path:
type: PathPrefix
value: /
backendRefs:
- name: app-v1
port: 8080
weight: 90
- name: app-v2
port: 8080
weight: 10
Like Ingress, Gateway uses container-native load balancing (NEGs) — the LB targets pod IPs directly, bypassing kube-proxy. Gateway API is the recommended path over the older Ingress for new workloads; use Ingress only for existing setups you are not ready to migrate.
Node Pools and GPU
# Add a GPU node pool
gcloud container node-pools create gpu-pool \
--cluster my-cluster \
--region us-central1 \
--machine-type n1-standard-4 \
--accelerator type=nvidia-tesla-t4,count=1 \
--num-nodes 0 \ # start at 0
--enable-autoscaling \
--min-nodes 0 \ # scale to zero when no GPU workloads
--max-nodes 4 \
--node-taints nvidia.com/gpu=present:NoSchedule
# Install NVIDIA drivers automatically (GKE manages this)
kubectl apply -f https://raw.githubusercontent.com/GoogleCloudPlatform/container-engine-accelerators/master/nvidia-driver-installer/cos/daemonset-preloaded.yaml
GKE Autopilot Limits
# Autopilot: minimum pod resource requests enforced
resources:
requests:
cpu: 250m # minimum 250m (Autopilot enforces)
memory: 512Mi # minimum 512Mi
limits:
cpu: 250m # requests == limits (Guaranteed QoS — Autopilot requirement)
memory: 512Mi
# No: DaemonSets, hostNetwork, privileged, hostPID
# No: node selectors for specific machine types
# Yes: GPU workloads (Autopilot provisions GPU nodes automatically)
On Autopilot, can you set a Pod's cpu limit higher than its cpu request, so it can burst above what it asked for when spare capacity is available?
GKE Upgrade Strategy
GKE decouples three separate questions: which version stream you're tracking (release channel), when Google is allowed to touch your nodes (maintenance window), and how a running node is actually replaced when an upgrade lands (surge upgrade). All three combine to determine how much control you have over disruption.
Release channels
# Check available versions
gcloud container get-server-config --region us-central1
# Enable auto-upgrade (recommended)
gcloud container node-pools update default-pool \
--cluster my-cluster \
--region us-central1 \
--enable-autoupgrade
# Use release channels (tracks stable/regular/rapid)
gcloud container clusters update my-cluster \
--region us-central1 \
--release-channel regular # regular = ~monthly, stable = ~quarterly
Node upgrade rollout
Whichever channel picks the target version, GKE never patches a node in place — it replaces it, node by node, with new capacity created before old capacity is torn down.
--max-surge-upgrade (default 1) — before touching anything running today.
--max-unavailable-upgrade (default 0) caps how many nodes can be draining at once alongside the surge.
During a GKE node upgrade, is the existing node patched to the new Kubernetes version in place, or replaced?
# Maintenance windows: only upgrade during specific hours
gcloud container clusters update my-cluster \
--maintenance-window-start 2024-01-15T02:00:00Z \
--maintenance-window-end 2024-01-15T06:00:00Z \
--maintenance-window-recurrence "FREQ=WEEKLY;BYDAY=SA,SU"