GCP Storage

Object storage, block storage, and file storage — GCP equivalents to S3, EBS, and EFS.

0/0 checks

Storage Services Map

Use case AWS GCP
Object storage S3 Cloud Storage (GCS)
Block storage (VM disk) EBS Persistent Disk / Hyperdisk
Shared file system (NFS) EFS Filestore
Local NVMe (ephemeral) Instance Store Local SSD
Transfer acceleration S3 Transfer Acceleration Storage Transfer Service
Archival S3 Glacier GCS Archive class
Auto-tiering S3 Intelligent-Tiering GCS Autoclass

Cloud Storage (GCS)

GCS is GCP's S3. Globally consistent, durable (11 nines), and multi-region by default when you want it.

Bucket Locations

Type Example Availability Cost
Region us-central1 Single region Lowest
Dual-region NAM4 (Iowa+S.Carolina) 2 regions, auto-replication Medium
Multi-region US, EU, ASIA 3+ regions Highest
graph TD
    classDef region fill:#3498db,stroke:#2471a3,color:#fff
    classDef dual fill:#e67e22,stroke:#ba6018,color:#fff
    classDef multi fill:#8e44ad,stroke:#6c3483,color:#fff

    subgraph REG["Region — e.g. us-central1"]
        R1["Single zone-redundant region<br/>lowest cost, lowest latency to one area<br/>no built-in cross-region copy"]:::region
    end

    subgraph DUAL["Dual-region — e.g. NAM4"]
        D1["Region A — us-central1"]:::dual
        D2["Region B — us-east1"]:::dual
        D1 -->|"synchronous auto-replication<br/>Turbo Replication SLA available"| D2
    end

    subgraph MULTI["Multi-region — e.g. US, EU, ASIA"]
        M1["Region 1"]:::multi
        M2["Region 2"]:::multi
        M3["Region 3+"]:::multi
        M1 --- M2
        M2 --- M3
        M1 --- M3
    end

Your app serves one geographic market and never needs to survive a full-region outage. Which bucket location type minimizes cost and latency, and what are you giving up compared to dual-region?

# Create a regional bucket
gsutil mb -l us-central1 gs://my-unique-bucket-name

# Multi-region (equivalent to S3 Cross-Region Replication, but built-in)
gsutil mb -l US gs://my-multi-region-bucket

# Uniform bucket-level access (recommended — disables per-object ACLs, use IAM only)
gsutil uniformbucketlevelaccess set on gs://my-bucket

Storage Classes

Class Min storage Retrieval cost Price/GB/mo Use case
Standard None None $0.020 Frequently accessed data
Nearline 30 days $0.01/GB $0.010 Monthly access (backups)
Coldline 90 days $0.02/GB $0.004 Quarterly access
Archive 365 days $0.05/GB $0.0012 Annual+ access (DR copies)

The table above is the number reference; the tabs below are the "which one do I actually pick" reference — same four classes, but framed around what happens if you get the minimum-duration commitment wrong.

No minimum storage duration, no retrieval fee. The default class for anything read or written more than about once a month: active application data, website assets, data a BigQuery/Dataflow job is currently chewing through. There's rarely a reason to pin an object here manually instead of letting Autoclass do it — Autoclass only costs you money if an object is genuinely hot the whole time.
gsutil cp -c STANDARD local-file.txt gs://my-bucket/
30-day minimum storage duration, small per-GB retrieval fee. Sized for data touched about once a month — nightly/weekly backups, warm DR secondaries. Move or delete an object before day 30 and you're still billed for the remainder of that 30-day window: the minimum duration is a commitment, not a suggestion.
gsutil rewrite -s nearline gs://my-bucket/large-file.parquet
90-day minimum, higher retrieval fee than Nearline. Quarterly-access data — compliance copies pulled once per audit cycle, analytics datasets that have gone cold but aren't dead. Same early-deletion math as Nearline, just a 90-day window instead of 30.
gsutil rewrite -s coldline gs://my-bucket/quarterly-report.parquet
365-day minimum, the highest retrieval fee — and still the cheapest class per GB stored. Annual-or-rarer access: long-term compliance retention, DR-of-DR copies, anything you sincerely hope to never read again. Deleting or moving one out inside a year still bills for the rest of the year.
gsutil -o "GSUtil:default_project_id=my-project" cp \
  -Z -c ARCHIVE \
  large-backup.tar.gz gs://my-backup-bucket/
# Upload to specific storage class
gsutil -o "GSUtil:default_project_id=my-project" cp \
  -Z -c ARCHIVE \
  large-backup.tar.gz gs://my-backup-bucket/

# Change class of existing object
gsutil rewrite -s nearline gs://my-bucket/large-file.parquet

You upload a file straight to Archive with -c ARCHIVE, then delete it 10 days later. Are you only billed for those 10 days?

Autoclass — Automatic Tiering

GCS Autoclass automatically moves objects between Standard → Nearline → Coldline → Archive based on access patterns. Equivalent to S3 Intelligent-Tiering.

stateDiagram-v2
    [*] --> Standard
    Standard --> Nearline: no access for 30 days
    Nearline --> Coldline: no access for 90 more days
    Coldline --> Archive: no access for 365 more days
    Nearline --> Standard: object read or written
    Coldline --> Standard: object read or written
    Archive --> Standard: object read or written
1. Object lands in Standard. Every new object starts Standard regardless of how it's eventually going to be accessed — Autoclass only downgrades based on observed behavior, it never guesses up front.
2. 30 days of no access → Nearline. Same day-count as Nearline's own minimum storage duration in the table above — Autoclass doesn't move an object earlier than the point where Nearline pricing would actually pay off.
3. 90 more days of no access → Coldline, then 365 more → Archive. The clock resets after each transition, not from the object's original upload date — it's "days since last touched," compounding down the tier list one step at a time.
4. Any read or write → back to Standard immediately. No waiting out a minimum duration first. Autoclass also waives the early-deletion fee for its own automatic transitions — the fee only applies when you manually move an object with gsutil rewrite before its minimum duration is up.
gsutil buckets update gs://my-bucket --autoclass

An object has been sitting untouched in Coldline for months under an Autoclass-enabled bucket. Someone reads it today. What class is it in tomorrow, and do you eat an early-deletion fee for leaving Coldline before its 90-day minimum was up?

Basic Operations

# Upload
gsutil cp local-file.txt gs://my-bucket/path/file.txt

# Upload directory (recursive)
gsutil cp -r ./local-dir/ gs://my-bucket/prefix/

# Sync (like aws s3 sync)
gsutil rsync -r ./local-dir/ gs://my-bucket/prefix/
gsutil rsync -r -d gs://my-bucket/prefix/ ./local-dir/   # -d deletes destination files not in source

# Download
gsutil cp gs://my-bucket/path/file.txt ./local-file.txt

# List
gsutil ls gs://my-bucket/
gsutil ls -l gs://my-bucket/    # long format with sizes

# Delete
gsutil rm gs://my-bucket/path/file.txt
gsutil rm -r gs://my-bucket/prefix/    # recursive

# Signed URL (pre-signed URL equivalent — time-limited access)
gsutil signurl -d 1h -m GET service-account-key.json gs://my-bucket/file.txt
# Better: use Workload Identity and generate via SDK, not key files

Access Control

# Grant read access to a service account (use IAM, not ACLs)
gcloud storage buckets add-iam-policy-binding gs://my-bucket \
  --member="serviceAccount:my-app@project.iam.gserviceaccount.com" \
  --role="roles/storage.objectViewer"

# Common storage IAM roles
# roles/storage.objectViewer      → read objects (not list bucket)
# roles/storage.objectUser        → read + write objects
# roles/storage.objectAdmin       → full object control
# roles/storage.admin             → bucket + object admin
# roles/storage.legacyBucketReader → list bucket (needed for gsutil ls)

Lifecycle Rules

{
  "rule": [
    {
      "action": {"type": "SetStorageClass", "storageClass": "NEARLINE"},
      "condition": {"age": 30, "matchesStorageClass": ["STANDARD"]}
    },
    {
      "action": {"type": "SetStorageClass", "storageClass": "COLDLINE"},
      "condition": {"age": 90}
    },
    {
      "action": {"type": "Delete"},
      "condition": {"age": 365}
    }
  ]
}
gsutil lifecycle set lifecycle.json gs://my-bucket

Rule 1 above only fires on objects that currently matchesStorageClass: ["STANDARD"]. Rule 2 has no matchesStorageClass condition at all — just {"age": 90}. What does omitting that filter actually do?

Versioning

# Enable versioning (equivalent to S3 versioning)
gsutil versioning set on gs://my-bucket

# List all versions
gsutil ls -a gs://my-bucket/file.txt

# Restore a version
gsutil cp gs://my-bucket/file.txt#1234567890 gs://my-bucket/file.txt

You run gsutil rm gs://my-bucket/file.txt on a bucket with versioning on. Is the data actually gone?


Persistent Disk (Block Storage)

Persistent Disk is GCP's EBS — network-attached block storage for GCE VMs.

Disk Types

Type Max IOPS Max throughput Price/GB/mo AWS analog
pd-standard 3,000 IOPS/TB 0.12 MB/s/GB $0.040 gp2 (old)
pd-balanced 3,000 IOPS/TB 0.28 MB/s/GB $0.100 gp3
pd-ssd 30,000 IOPS 0.48 MB/s/GB $0.170 io1/io2
pd-extreme 120,000 IOPS Custom $0.220 io2 Block Express
hyperdisk-balanced 160,000 IOPS Configurable $0.120 io2 Express
hyperdisk-throughput 3,000 IOPS 2,400 MB/s $0.080 Throughput-optimized
# Create and attach a disk
gcloud compute disks create my-data-disk \
  --zone=us-central1-a \
  --size=200GB \
  --type=pd-ssd

gcloud compute instances attach-disk my-vm \
  --disk=my-data-disk \
  --device-name=data \
  --zone=us-central1-a

# Inside VM: format and mount
sudo mkfs.ext4 -m 0 -E lazy_itable_init=0,lazy_journal_init=0,discard /dev/disk/by-id/google-data
sudo mkdir -p /mnt/data
sudo mount /dev/disk/by-id/google-data /mnt/data

# Resize a disk online (no reboot needed — unlike AWS)
gcloud compute disks resize my-data-disk \
  --size=400GB \
  --zone=us-central1-a
# Then grow filesystem online:
sudo resize2fs /dev/disk/by-id/google-data

You run gcloud compute disks resize to grow a disk from 200GB to 400GB while the VM keeps serving traffic. Can the application immediately use the extra 200GB?

Multi-Reader Disks

A PD disk can be attached to multiple VMs in read-only mode. Useful for shared datasets (ML model weights, reference data).

# Attach same disk to multiple VMs in read-only mode
gcloud compute instances attach-disk vm-1 --disk=my-shared-disk --mode=ro --zone=us-central1-a
gcloud compute instances attach-disk vm-2 --disk=my-shared-disk --mode=ro --zone=us-central1-a
gcloud compute instances attach-disk vm-3 --disk=my-shared-disk --mode=ro --zone=us-central1-a
graph TD
    classDef disk fill:#4285f4,stroke:#2a56c6,color:#fff
    classDef vm fill:#e67e22,stroke:#ba6018,color:#fff
    classDef awsdisk fill:#ff9900,stroke:#cc7a00,color:#fff
    classDef awsvm fill:#232f3e,stroke:#0f1721,color:#fff
    classDef bad fill:#c0392b,stroke:#8e2418,color:#fff

    subgraph GCP["GCP — any Persistent Disk type, read-only"]
        PD["my-shared-disk<br/>pd-standard / pd-balanced / pd-ssd — all supported"]:::disk
        VM1["vm-1 — mode=ro"]:::vm
        VM2["vm-2 — mode=ro"]:::vm
        VM3["vm-3 — mode=ro"]:::vm
        PD --> VM1
        PD --> VM2
        PD --> VM3
    end

    subgraph AWS["AWS — EBS Multi-Attach"]
        EBS["io1 / io2 volume only"]:::awsdisk
        AVM1["instance-1<br/>cluster-aware app required"]:::awsvm
        AVM2["instance-2<br/>cluster-aware app required"]:::awsvm
        BLOCKED["gp2 / gp3 / st1 volumes<br/>Multi-Attach not offered"]:::bad
        EBS --> AVM1
        EBS --> AVM2
        EBS -.->|"unsupported on these types"| BLOCKED
    end

AWS EBS multi-attach is only supported on io1/io2 and only for cluster-aware applications. GCP PD read-only multi-attach works for any disk type.

What does GCP trade off to let PD multi-attach work on any disk type, where AWS restricts EBS Multi-Attach to io1/io2 and cluster-aware apps?

Disk Snapshots

# Manual snapshot (= EBS snapshot)
gcloud compute disks snapshot my-data-disk \
  --snapshot-names=my-data-disk-snap-$(date +%Y%m%d) \
  --zone=us-central1-a

# Restore from snapshot
gcloud compute disks create restored-disk \
  --source-snapshot=my-data-disk-snap-20240115 \
  --zone=us-central1-a

# Scheduled snapshots (= Data Lifecycle Manager in AWS)
gcloud compute resource-policies create snapshot-schedule daily-backup \
  --region=us-central1 \
  --max-retention-days=7 \
  --on-source-disk-delete=keep-auto-snapshots \
  --daily-schedule \
  --start-time=04:00

gcloud compute disks add-resource-policies my-data-disk \
  --resource-policies=daily-backup \
  --zone=us-central1-a

Local SSD

Local NVMe drives physically attached to the server. Much faster than Persistent Disk, but ephemeral — data is lost on VM stop/restart. Like AWS instance store.

# Attach local SSD at instance creation (can't attach after)
gcloud compute instances create my-vm \
  --machine-type=n2-standard-8 \
  --local-ssd=interface=nvme    # 375 GB per SSD

# Multiple local SSDs (can be striped for more throughput)
gcloud compute instances create my-vm \
  --machine-type=n2-standard-16 \
  --local-ssd=interface=nvme \
  --local-ssd=interface=nvme    # 750 GB total

Use cases: temp files, shuffle space for batch jobs, buffer layers on top of Persistent Disk.

A batch job writes its shuffle data to Local SSD for the throughput. Midway through, the VM is stopped and restarted (not just rebooted from inside the guest OS). What happens to that shuffle data?


Filestore (Managed NFS)

Filestore is GCP's EFS — managed NFS server for shared file access across multiple VMs or GKE pods.

graph TD
    classDef compute fill:#e67e22,stroke:#ba6018,color:#fff
    classDef gke fill:#16a085,stroke:#117a65,color:#fff
    classDef fs fill:#4285f4,stroke:#2a56c6,color:#fff
    classDef csi fill:#8e44ad,stroke:#6c3483,color:#fff

    subgraph CLIENTS["Compute clients"]
        VM1["GCE VM 1<br/>standard NFS client"]:::compute
        VM2["GCE VM 2<br/>standard NFS client"]:::compute
        subgraph GKENS["GKE cluster"]
            POD["Pods"]:::gke
            PVC["PVC — accessModes: ReadWriteMany"]:::csi
            CSI["Filestore CSI driver<br/>pre-installed on GKE"]:::csi
            POD --> PVC --> CSI
        end
    end

    FS["Filestore instance<br/>10.0.0.100:/vol<br/>NFSv3, single mount target"]:::fs

    VM1 -->|"sudo mount 10.0.0.100:/vol"| FS
    VM2 -->|"sudo mount 10.0.0.100:/vol"| FS
    CSI -->|"provisions & mounts volume"| FS

Filestore Tiers

Tier Capacity IOPS Use case AWS analog
Basic HDD 1-63.9 TB 600/TB Dev/test EFS IA
Basic SSD 2.5-63.9 TB 30,000 Production web/content EFS General
Enterprise 1-10 TB 120,000 Databases, high-perf EFS Max I/O
Zonal 1-9.75 TB 80,000 Single-zone, lower cost EFS Standard
# Create Filestore instance
gcloud filestore instances create my-nfs \
  --zone=us-central1-a \
  --tier=BASIC_SSD \
  --file-share=name=vol,capacity=2.5TB \
  --network=name=my-vpc

# Mount on VM
sudo apt-get install nfs-common
sudo mkdir /mnt/shared
sudo mount 10.0.0.100:/vol /mnt/shared

# Mount in GKE (via CSI driver)
# StorageClass is pre-installed on GKE
# PVC for Filestore in GKE
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: my-nfs-pvc
spec:
  accessModes:
    - ReadWriteMany     # multiple pods can mount simultaneously
  storageClassName: standard-rwx   # uses Filestore CSI driver
  resources:
    requests:
      storage: 2560Gi

You're standing up a shared volume for a dev/test environment where cost matters far more than throughput. Which tier fits, and why would picking it for a production database be a mistake?


Storage Transfer Service

Move data into GCS from S3, Azure Blob, HTTP, on-prem. Equivalent to AWS DataSync or S3 Transfer Acceleration (for migration).

sequenceDiagram
    participant U as You — gcloud transfer jobs create
    participant STS as Storage Transfer Service
    participant SRC as Source — S3 / Azure Blob / HTTP / on-prem
    participant DST as GCS bucket

    U->>STS: Define source, destination, and schedule
    STS->>SRC: Authenticate using source credentials or agent pool
    STS->>SRC: List objects at the source
    STS->>DST: List objects already at the destination
    STS->>STS: Diff the two listings
    STS->>DST: Copy only the missing or changed objects
    Note over STS,DST: Job then sleeps until the next scheduled run
    loop every schedule-repeats-every interval
        STS->>SRC: Re-list and re-diff
        STS->>DST: Copy any new or changed objects only
    end
1. Define the job. Source (S3, Azure Blob, HTTP, or on-prem), destination GCS bucket, and optionally a repeat schedule like --schedule-repeats-every=24h.
2. Authenticate to the source. This is a pull, not a push — GCP reaches out to the source, so the job needs read credentials for it (--source-creds-file for S3) or, for an on-prem source, an agent pool it can reach over the network (--source-agent-pool).
3. List and diff. Before copying a single byte, the job lists what's already sitting in the destination bucket and compares it against the source listing.
4. Copy only what's different. Just the new or changed objects transfer — not a full re-copy of everything, every time.
5. Repeat on schedule. If a repeat interval was set, the whole list → diff → copy cycle runs again automatically without anyone re-triggering it.
# Create a transfer job from S3 to GCS
gcloud transfer jobs create \
  --source-agent-pool="" \
  --source-creds-file=aws-creds.json \
  --source=s3://my-aws-bucket/ \
  --destination=gs://my-gcp-bucket/ \
  --schedule-repeats-every=24h

The job above points at --source-creds-file=aws-creds.json rather than any GCP-side setting for reading from S3. What does that tell you about which side initiates the transfer, and what you need to configure where?


GCS vs S3 — Key Differences

GCP Cloud Storage AWS S3
Consistency Strong consistency (always) Strong consistency (since 2020)
Storage classes Standard/Nearline/Coldline/Archive Standard/IA/Glacier/Deep Archive
Auto-tiering Autoclass Intelligent-Tiering
Multi-region US/EU/ASIA built-in Cross-Region Replication (separate config)
Access control Uniform (IAM) or fine-grained (ACLs) Bucket policies + ACLs
Signed URLs gsutil signurl or client library aws s3 presign
Event notifications Pub/Sub, Eventarc S3 Events → SNS/SQS/Lambda
Versioning Yes Yes
Lifecycle rules Yes Yes
Replication Turbo Replication (dual-region) CRR/SRR
CLI gsutil or gcloud storage aws s3

A team designs read-after-write workarounds for S3 because "object storage is eventually consistent." Per the table, are they right today, and does GCS need the same workaround?

# New gcloud storage CLI (faster than gsutil for large transfers)
gcloud storage cp local-file.txt gs://my-bucket/
gcloud storage ls gs://my-bucket/
gcloud storage rm gs://my-bucket/file.txt