GCP Observability — Cloud Monitoring, Logging, Trace, Audit Logs
Observability Service Map
| Pillar | AWS | GCP |
|---|---|---|
| Metrics | CloudWatch Metrics | Cloud Monitoring |
| Logs | CloudWatch Logs | Cloud Logging |
| Traces | X-Ray | Cloud Trace |
| Dashboards | CloudWatch Dashboards | Cloud Monitoring Dashboards |
| Alerts | CloudWatch Alarms | Cloud Monitoring Alerting |
| API audit trail | CloudTrail | Cloud Audit Logs |
| Uptime checks | Route 53 Health Checks | Cloud Monitoring Uptime Checks |
| Error tracking | — | Error Reporting |
| Profiling | — | Cloud Profiler |
GCP advantage: GKE, Cloud Run, Cloud SQL, GCE all emit structured logs and metrics automatically. In AWS, you often need CloudWatch Agent configuration before metrics appear.
You spin up a new GKE deployment and a new Cloud Run service. Before either has served a single request, do their CPU and memory metrics already exist in Cloud Monitoring?
Cloud Monitoring
Cloud Monitoring collects metrics from all GCP services automatically. You don't configure agents for managed services.
Metric Types
graph TD
classDef auto fill:#2980b9,stroke:#1f618d,color:#fff
classDef custom fill:#8e44ad,stroke:#6c3483,color:#fff
classDef store fill:#27ae60,stroke:#1e8449,color:#fff
subgraph AUTO["System metrics — auto-collected, zero configuration"]
M1["compute.googleapis.com/instance/cpu/utilization<br/>GCE VM CPU"]:::auto
M2["kubernetes.io/container/memory/used_bytes<br/>GKE container memory"]:::auto
M3["run.googleapis.com/request_count<br/>Cloud Run request volume"]:::auto
M4["cloudsql.googleapis.com/database/cpu/utilization<br/>Cloud SQL CPU"]:::auto
end
subgraph CUSTOM["Custom metrics — you emit them explicitly"]
C1["custom.googleapis.com/myapp/orders_processed<br/>business counter"]:::custom
C2["custom.googleapis.com/myapp/queue_depth<br/>application gauge"]:::custom
end
AUTO -.->|"no agent, no SDK call needed —<br/>the managed service reports for you"| MON["Cloud Monitoring<br/>time-series store"]:::store
CUSTOM -.->|"MetricServiceClient.create_time_series()<br/>your code decides when and what"| MON
MetricServiceClient.create_time_series(). Cloud Monitoring stores and graphs it exactly like a system metric once it arrives, but nothing shows up until your code calls the API.
Emit Custom Metrics
from google.cloud import monitoring_v3
import time
client = monitoring_v3.MetricServiceClient()
project_name = f"projects/my-project"
series = monitoring_v3.TimeSeries()
series.metric.type = "custom.googleapis.com/myapp/orders_processed"
series.metric.labels["environment"] = "production"
series.resource.type = "global"
now = time.time()
interval = monitoring_v3.TimeInterval({
"end_time": {"seconds": int(now), "nanos": 0}
})
point = monitoring_v3.Point({
"interval": interval,
"value": {"int64_value": 42}
})
series.points = [point]
client.create_time_series(name=project_name, time_series=[series])
Alerting Policies
# Create an alert when CPU > 80% for 5 minutes
gcloud alpha monitoring policies create \
--notification-channels=projects/my-project/notificationChannels/12345 \
--display-name="High CPU Alert" \
--condition-display-name="CPU > 80%" \
--condition-filter='resource.type="gce_instance" AND metric.type="compute.googleapis.com/instance/cpu/utilization"' \
--condition-threshold-value=0.8 \
--condition-threshold-comparison=COMPARISON_GT \
--condition-threshold-duration=300s
Or use Terraform (recommended for production):
resource "google_monitoring_alert_policy" "cpu_alert" {
display_name = "High CPU Alert"
combiner = "OR"
conditions {
display_name = "CPU utilization > 80%"
condition_threshold {
filter = "resource.type=\"gce_instance\" AND metric.type=\"compute.googleapis.com/instance/cpu/utilization\""
duration = "300s"
comparison = "COMPARISON_GT"
threshold_value = 0.8
}
}
notification_channels = [google_monitoring_notification_channel.email.id]
}
graph LR
classDef metric fill:#2980b9,stroke:#1f618d,color:#fff
classDef eval fill:#e67e22,stroke:#ba6018,color:#fff
classDef incident fill:#e74c3c,stroke:#c0392b,color:#fff
classDef notify fill:#27ae60,stroke:#1e8449,color:#fff
TS["Time series<br/>cpu/utilization samples"]:::metric --> COND{"condition_threshold met?<br/>value > 0.8 for the full<br/>duration: 300s window"}:::eval
COND -->|"No, or it dips<br/>back down early"| TS
COND -->|"Yes — sustained for<br/>the entire duration"| INC["Incident opened"]:::incident
INC --> NOTIF["Notification channels<br/>email / Slack / PagerDuty"]:::notify
INC -.->|"condition clears<br/>and stays clear"| RESOLVE["Incident auto-resolved"]:::notify
CPU spikes to 95% for 30 seconds, then drops back to 40%. The alert policy above requires CPU > 80% for a duration of 300s. Does it fire?
duration field means the condition has to stay true for the entire window — a 30-second spike that then drops back below the threshold never accumulates 300 continuous seconds above 80%, so the timer resets and no incident opens. This is what keeps alerting policies from firing on brief, self-correcting blips.Uptime Checks (= Route 53 Health Checks)
gcloud monitoring uptime create my-service-uptime \
--display-name="My API Uptime" \
--resource-type=uptime-url \
--hostname=api.example.com \
--path=/health \
--port=443 \
--use-ssl \
--check-interval=60s
Cloud Logging
All GCP services write structured logs automatically. Cloud Run, GKE containers, Cloud Functions, Cloud SQL — everything logs to Cloud Logging without agent setup.
graph LR
classDef app fill:#34495e,stroke:#212f3c,color:#fff
classDef ingest fill:#2980b9,stroke:#1f618d,color:#fff
classDef router fill:#8e44ad,stroke:#6c3483,color:#fff
classDef sink fill:#e67e22,stroke:#ba6018,color:#fff
classDef view fill:#27ae60,stroke:#1e8449,color:#fff
APP["Cloud Run / GKE / Cloud Functions<br/>writes structured JSON to stdout"]:::app --> AGENT["Built-in logging agent<br/>no install, no config"]:::ingest
AGENT --> ROUTER["Log Router<br/>evaluates every sink's filter<br/>against each incoming entry"]:::router
ROUTER --> EXPLORER["Logs Explorer<br/>LQL queries, kept per<br/>the retention table below"]:::view
ROUTER --> METRIC["Log-based metric<br/>e.g. error-rate counter"]:::view
ROUTER -.->|"filter matches"| BQ["BigQuery sink<br/>SQL analytics"]:::sink
ROUTER -.->|"filter matches"| GCS["GCS sink<br/>long-term archival"]:::sink
ROUTER -.->|"filter matches"| PS["Pub/Sub sink<br/>stream to a third party"]:::sink
Log Levels and Structure
import logging
from google.cloud import logging as cloud_logging
# In Cloud Run / GKE: just write structured JSON to stdout
import json, sys
def log(severity, message, **kwargs):
entry = {
"severity": severity,
"message": message,
"component": "order-service",
**kwargs
}
print(json.dumps(entry), flush=True)
log("INFO", "Order placed", order_id="123", amount=99.99)
log("ERROR", "Payment failed", order_id="123", error="card_declined")
Cloud Logging automatically indexes structured JSON fields — you can query by jsonPayload.order_id directly.
You log a structured entry containing jsonPayload.order_id = "123". Do you need to configure a schema or build an index before you can filter on jsonPayload.order_id in a query?
Log Queries (Cloud Logging Query Language)
# Query logs via CLI
gcloud logging read 'resource.type="k8s_container" AND severity=ERROR' \
--limit=50 \
--freshness=1h \
--format=json
# Common filters
gcloud logging read 'resource.type="cloud_run_revision"
AND resource.labels.service_name="my-api"
AND severity>=WARNING
AND timestamp>="2024-01-15T00:00:00Z"' \
--limit=100
# Logging Query Language (LQL) — used in console and API
# Filter by log level
severity >= WARNING
# Filter by resource
resource.type = "k8s_container"
resource.labels.namespace_name = "production"
# Filter by log field (structured JSON)
jsonPayload.order_id = "123"
jsonPayload.http_status >= 500
# Full text search
textPayload: "connection refused"
# Time range
timestamp >= "2024-01-15T00:00:00Z" AND timestamp <= "2024-01-15T01:00:00Z"
# Combine
resource.type="cloud_run_revision"
AND jsonPayload.http_status>=500
AND severity=ERROR
Log-Based Metrics
Create metrics from log patterns — like CloudWatch Metric Filters:
# Count ERROR logs per service
gcloud logging metrics create error-rate \
--description="Errors per service" \
--log-filter='severity=ERROR AND resource.type="cloud_run_revision"' \
--value-extractor='EXTRACT(jsonPayload.latency_ms)' # optional: extract numeric value
Log Sinks — Export to BigQuery / GCS / Pub/Sub
# Export all ERROR logs to BigQuery for analysis
gcloud logging sinks create errors-to-bigquery \
bigquery.googleapis.com/projects/my-project/datasets/logs \
--log-filter='severity >= ERROR'
# Export all logs to GCS (long-term archival)
gcloud logging sinks create all-logs-to-gcs \
storage.googleapis.com/my-logs-bucket \
--log-filter='' \
--include-children # include logs from child resources
Log Retention
| Log type | Default retention | Configurable |
|---|---|---|
| Admin Activity | 400 days | No (always kept) |
| Data Access | 30 days | Yes (1–3650 days) |
| System Event | 400 days | No |
| User-written | 30 days | Yes |
A compliance team wants Admin Activity logs kept for only 90 days instead of the default, to cut storage cost. Can they configure that?
Cloud Trace
Cloud Trace = AWS X-Ray. Distributed tracing across services.
For GKE and Cloud Run, traces can be auto-collected via OpenTelemetry collector.
Auto-instrumentation with OpenTelemetry
# requirements.txt
# opentelemetry-api
# opentelemetry-sdk
# opentelemetry-exporter-gcp-trace
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.exporter.cloud_trace import CloudTraceSpanExporter
from opentelemetry.sdk.trace.export import BatchSpanProcessor
# Setup
provider = TracerProvider()
exporter = CloudTraceSpanExporter(project_id="my-project")
provider.add_span_processor(BatchSpanProcessor(exporter))
trace.set_tracer_provider(provider)
tracer = trace.get_tracer(__name__)
# Instrument your code
def process_order(order_id: str):
with tracer.start_as_current_span("process_order") as span:
span.set_attribute("order.id", order_id)
with tracer.start_as_current_span("validate_payment"):
validate_payment(order_id)
with tracer.start_as_current_span("update_inventory"):
update_inventory(order_id)
Trace Context Propagation
Every span in the process_order example above shares one thing: a single trace ID. That's not automatic magic inside one process — it's a context object passed down through every nested start_as_current_span call. The moment a request crosses a network boundary into another service, that context has to be serialized into a header and picked back up on the other side, or the two services show up as two disconnected traces instead of one.
sequenceDiagram
participant Client
participant OrderSvc as order-service
participant PaySvc as payment-service
participant CT as Cloud Trace
Client->>OrderSvc: POST /orders, no trace header yet
OrderSvc->>OrderSvc: start_as_current_span(process_order), mint new trace ID
OrderSvc->>OrderSvc: start_as_current_span(validate_payment), child span
OrderSvc->>PaySvc: HTTP call, inject traceparent header from active context
PaySvc->>PaySvc: extract traceparent, start child span in the same trace
PaySvc-->>OrderSvc: response
OrderSvc->>OrderSvc: end validate_payment span
OrderSvc-->>Client: order confirmed
OrderSvc->>CT: BatchSpanProcessor exports its spans, same trace ID
PaySvc->>CT: exports its own span, same trace ID
Note over CT: Console groups every span sharing<br/>one trace ID into a single waterfall
traceparent header, so it mints a brand-new trace ID and starts the root span — this is exactly the process_order span in the code above.
with tracer.start_as_current_span(...) block inside process_order — validate_payment, update_inventory — inherits the active context automatically, so every child span carries the same trace ID as its parent.
traceparent header — without this step, payment-service has no way to know it's part of the same request.
BatchSpanProcessor ships its own spans to Cloud Trace on its own schedule — Cloud Trace has no idea they're related until it groups every span sharing the same trace ID into one waterfall in the console.
order-service calls payment-service over HTTP, but the outgoing request never gets a traceparent header injected. What shows up in Cloud Trace?
Auto-instrumentation in GKE (OpenTelemetry Operator)
# Annotate a deployment for auto-instrumentation (no code changes)
apiVersion: apps/v1
kind: Deployment
metadata:
name: my-service
annotations:
instrumentation.opentelemetry.io/inject-python: "true"
spec:
template:
metadata:
labels:
app: my-service
Cloud Audit Logs
Every GCP API call is logged in Audit Logs. Equivalent to CloudTrail — but mandatory for Admin Activity (cannot be disabled).
Log Types
| Type | What it captures | Default enabled |
|---|---|---|
| Admin Activity | Create/delete/modify resources | Always on |
| Data Access | Read data, get metadata | Off by default (verbose + costly) |
| System Event | Google-initiated (live migrations, auto-repairs) | Always on |
| Policy Denied | Requests denied by VPC Service Controls | Always on |
graph TD
classDef always fill:#27ae60,stroke:#1e8449,color:#fff
classDef optin fill:#f39c12,stroke:#ba6018,color:#fff
classDef q fill:#34495e,stroke:#212f3c,color:#fff
CALL["Any GCP API call"]:::q --> Q1{"Did Google trigger it,<br/>not a user or service account?"}:::q
Q1 -->|Yes| SYS["System Event<br/>always on, 400 days"]:::always
Q1 -->|No| Q2{"Was it blocked by<br/>VPC Service Controls?"}:::q
Q2 -->|Yes| POL["Policy Denied<br/>always on"]:::always
Q2 -->|No| Q3{"Does it create, delete,<br/>or modify a resource?"}:::q
Q3 -->|Yes| ADMIN["Admin Activity<br/>always on, 400 days,<br/>cannot be disabled"]:::always
Q3 -->|"No — it's a read of<br/>data or metadata"| DATA["Data Access<br/>off by default,<br/>30 days if enabled"]:::optin
A team wants to audit who is reading objects out of a specific Cloud Storage bucket. Do Admin Activity logs already cover that?
# Query audit logs for all GCS deletions
gcloud logging read \
'protoPayload.serviceName="storage.googleapis.com"
AND protoPayload.methodName="storage.objects.delete"' \
--freshness=24h
# Query for who modified IAM policies
gcloud logging read \
'protoPayload.methodName="SetIamPolicy"
AND protoPayload.serviceName="iam.googleapis.com"' \
--freshness=7d
Enable Data Access Logs
# Enable for Cloud Storage (logs all read/write access)
gcloud projects get-iam-policy my-project > policy.yaml
# Add to policy.yaml:
# auditConfigs:
# - auditLogConfigs:
# - logType: DATA_READ
# - logType: DATA_WRITE
# service: storage.googleapis.com
gcloud projects set-iam-policy my-project policy.yaml
Error Reporting
Automatically groups exceptions from Cloud Run, GKE, App Engine, and Cloud Functions. No setup needed — just let exceptions bubble with a stack trace.
# View errors
gcloud error-reporting events list --service=my-api --version=v2
In the console: Error Reporting shows count, first/last seen, affected users, and the stack trace grouped by error signature.
Your Cloud Run service throws an unhandled exception with a stack trace, written to stderr. Do you need to install or configure anything for it to show up in Error Reporting?
Cloud Profiler
Continuous CPU and memory profiling in production. No sampling gaps, minimal overhead (<1%). Equivalent to AWS CodeGuru Profiler.
# Add to your main.py
import googlecloudprofiler
googlecloudprofiler.start(
service="my-api",
service_version="1.0.0",
project_id="my-project"
)
Shows flame graphs in the GCP console — where CPU time is actually spent in production.
A teammate says continuous profiling in production is too risky — "won't it slow everything down under load?" What does Cloud Profiler's own design say about that tradeoff?
Monitoring Stack Summary
| Need | Tool |
|---|---|
| VM / GKE / Cloud Run metrics | Cloud Monitoring (automatic) |
| Custom business metrics | Cloud Monitoring custom metrics |
| Application logs | Cloud Logging (write structured JSON to stdout) |
| Log queries and alerts | Cloud Logging + Log-based metrics |
| Distributed tracing | Cloud Trace + OpenTelemetry |
| Audit trail (who did what) | Cloud Audit Logs |
| Exception tracking | Error Reporting |
| Production CPU profiling | Cloud Profiler |
| Long-term log archival | Log Sink → GCS |
| Log analytics with SQL | Log Sink → BigQuery |
| Third-party (Grafana, etc.) | Metrics via Cloud Monitoring API |
GCP vs AWS Observability
| GCP | AWS | |
|---|---|---|
| Auto-instrumentation | GKE + Cloud Run log/metric automatically | Requires CloudWatch agent for EC2 |
| Log format | Structured JSON indexed immediately | Raw text (Insights adds parsing) |
| Trace format | CloudEvents / OTel (open standard) | X-Ray format (proprietary) |
| Audit log retention | Admin Activity: 400 days (permanent) | CloudTrail: configurable, paid |
| Uptime checks | Free (50 checks) | Route 53 health checks ($0.50/check/mo) |
| Error grouping | Error Reporting (automatic) | CloudWatch (manual) |