TrademarkTrademark
Features
Documentation
  1. Learning Center
  2. Terraform Troubleshooting, Optimization and Error Resolution

Article · part of a guide

AWS Provider Memory Explosion: The v4.67.0+ Survival Guide

Terraform AWS provider v4.67.0 triggers a memory explosion; learn the cause, impact, and fast fixes in our practical survival guide.

AWS Provider Memory Explosion: The v4.67.0+ Survival Guide

Key takeaways

  1. AWS provider v4.67.0 added QuickSight resources with deeply nested schemas, and in one public benchmark max memory rose from 558MB on v4.66.1 to 729MB on v4.67.0; the same benchmark shows v5.20.0 down to 194MB.
  2. Pinning the provider to v4.66.1 stops the memory growth but costs you 9+ months of AWS features and bug fixes.
  3. Provider plugin caching speeds up init by skipping repeat provider downloads, while reducing parallelism trades execution speed for lower peak memory.
  4. Splitting monolithic configurations and trimming large state files reduces memory per run.
  5. At 200+ resources or many accounts, managed platforms like Scalr pre-cache providers and scale memory automatically rather than requiring manual optimization.

The Memory Crisis: What Actually Happened

AWS Provider v4.67.0 introduced QuickSight resources with large nested schemas. Terraform loads every provider schema during initialization, whether or not you use those resources, so the size of the schema lands on every run. Here's what that looks like in one public benchmark (max RSS, 17 regions, Terraform 1.4.5). The same benchmark shows v5.20.0 on Terraform 1.6.0 at 194MB, so check a current provider version before tuning around the problem:

# Memory usage comparison
v4.66.1: 558MB
v4.67.0: 729MB (+31%)
v5.1.0:  1,102MB (+97%)

In the GitHub issue that first reported the regression, a setup refreshing data sources across 17 regions saw max memory go from about 2.6GB on v4.66.1 to 3.6GB on v4.67.0. Each provider alias runs as its own plugin process, so multi-region deployments are hit particularly hard.

Immediate Fixes for Production Environments

Version Pinning Strategy

The first thing to do is pin your provider version while you work on the other fixes.

terraform {
  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = "4.66.1" # Last version before memory explosion
    }
  }
}

The cost is that you're now missing 9+ months of AWS features and bug fixes, which hurts teams that need the latest AWS services.

Docker Memory Configuration

Production Docker deployments need significant memory headroom:

version: '3.8'
services:
  terraform:
    image: hashicorp/terraform:1.6.0
    deploy:
      resources:
        limits:
          memory: 8G
        reservations:
          memory: 4G
    environment:
      - TF_PLUGIN_CACHE_DIR=/opt/terraform/plugin-cache
      - TF_CLI_ARGS_plan=-parallelism=3
    volumes:
      - terraform-cache:/opt/terraform/plugin-cache

This configuration can prevent OOM kills, but you're paying for a bigger container than the same workload needed before the regression.

GitHub Actions Optimization

GitHub Actions runners have fixed memory limits, so the workflow has to keep peak memory down:

name: Terraform Deploy
on: [push]
 
jobs:
  terraform:
    runs-on: ubuntu-latest
    env:
      TF_PLUGIN_CACHE_DIR: ${{ github.workspace }}/.terraform.d/plugin-cache
      TF_CLI_ARGS_plan: -parallelism=2
      
    steps:
    - uses: actions/checkout@v4
    
    - name: Cache Terraform providers
      uses: actions/cache@v3
      with:
        path: ${{ github.workspace }}/.terraform.d/plugin-cache
        key: terraform-providers-${{ hashFiles('**/.terraform.lock.hcl') }}
        
    - name: Terraform Init
      run: |
        # Monitor memory during init
        free -m
        terraform init
        free -m

Reducing parallelism from the default of 10 to 2 lowers peak memory but makes runs slower, so you're trading speed for stability.

Provider Plugin Caching

Provider caching helps but isn't a silver bullet:

# Setup provider cache
mkdir -p ~/.terraform.d/plugin-cache
 
cat > ~/.terraformrc << 'EOF'
plugin_cache_dir = "$HOME/.terraform.d/plugin-cache"
plugin_cache_may_break_dependency_lock_file = true
EOF

Caching saves the provider download on each init, which makes init faster, but it doesn't shrink the schema the provider loads into memory, so expect base memory usage to stay about the same.

Memory Profiling and Diagnostics

Knowing where the memory goes tells you which fix to try first:

# Enable detailed logging (RPC/plugin/graph diagnostics; Terraform does not log memory data)
export TF_LOG=DEBUG
 
# Observe the process's memory while a plan runs
/usr/bin/time -v terraform plan   # see "Maximum resident set size"
 
# Or watch system memory during runs
watch -n 1 'free -m | grep Mem'

AWS provider maintainers attribute much of its memory use to deeply nested schemas allocated when the provider starts, so provider schema loading is usually the first place to look.

Long-term Solutions and Architecture Changes

Infrastructure Segmentation

Breaking monolithic configurations reduces memory per run:

infrastructure/
├── networking/        # 200MB memory
├── compute/          # 300MB memory
├── data/            # 250MB memory
└── monitoring/      # 150MB memory

Instead of one large process, you get four smaller ones, though now you're managing state dependencies by hand.

State File Optimization

Large state files compound the problem:

# Check state size
terraform state pull | wc -c
 
# Remove unused resources
terraform state list | grep -E "null_resource" | \
  xargs -I {} terraform state rm {}
 
# Compact state
terraform state pull | jq -c . > compact.json
terraform state push compact.json

Compacting only strips whitespace from the JSON, so it shrinks the file but does little for memory; removing resources you no longer manage is what actually makes state smaller.

When to Consider Managed Platforms

Look at where the engineering time is going. By this point you're spending it on:

  • Memory optimization instead of infrastructure
  • Workarounds for provider limitations
  • CI/CD pipeline complexity
  • State file management overhead

A managed platform takes that work off your team. Scalr, for instance, runs Terraform in optimized environments with:

  • Pre-cached providers across workspaces
  • A set memory limit per run container (2 GB by default), which you can raise on self-hosted agents
  • State management without manual optimization
  • No need to maintain Docker configurations or GitHub Actions workflows

For the cost comparison, add up:

  • Engineering hours spent on memory optimization
  • Increased infrastructure costs (larger runners)
  • Pipeline complexity maintenance
  • Risk of OOM failures in production

To run that comparison against a managed option, Scalr's current plans are on the pricing page.

Summary and Recommendations

A rough decision matrix:

Scenario Memory Usage Complexity Recommendation
<50 resources, single region 2GB Low Self-managed with caching
50-200 resources, multi-region 4-8GB Medium Consider managed platform
200+ resources, many accounts 8GB+ High Managed platform recommended
Enterprise scale 16GB+ Very High Managed platform essential

Immediate Actions

  1. Upgrade to Terraform 1.6.0+ - Adds global provider schema caching, which cuts memory when a configuration has many instances of the same provider
  2. Implement provider caching - Faster init, fewer downloads
  3. Reduce parallelism - Trade speed for stability
  4. Monitor memory usage - Know your baseline

Strategic Decisions

If you keep hitting memory limits, it's worth asking whether tuning Terraform memory is where your engineering time should go. Managed platforms like Scalr take that tuning work off your plate.

The AWS Provider memory issue isn't going away soon. QuickSight resources are here to stay, and AWS keeps adding complex services, so plan for it: either invest in DIY optimizations or move to a platform built to handle this kind of load.

Frequently asked questions

Why does the Terraform AWS provider use so much memory after v4.67.0?

AWS provider v4.67.0 introduced QuickSight resources with large nested schemas, and Terraform loads every provider schema during initialization whether or not you use those resources. In one public benchmark, max memory went from 558MB on v4.66.1 to 729MB on v4.67.0 and 1,102MB on v5.1.0, then fell to 194MB on v5.20.0 with Terraform 1.6.0. Each provider alias also runs as its own plugin process, so multi-region setups are hit hardest.

How do I reduce Terraform AWS provider memory usage?

Upgrade first: in a public benchmark, AWS provider v5.20.0 used far less memory than v5.1.0. Then enable provider plugin caching so init stops re-downloading providers, and reduce plan parallelism, which lowers peak memory at the cost of slower runs. You can also split monolithic configurations into smaller ones.

Should I pin the AWS provider to v4.66.1 to avoid the memory issue?

Pinning to v4.66.1 stops the memory growth immediately, so it works as a first response while you apply other fixes. The cost is that you miss out on more than nine months of AWS features and bug fixes, which isn't ideal for teams that need newer AWS services.

When is a managed Terraform platform worth it for memory problems?

At around 200+ resources or many AWS accounts, self-managed optimization gets expensive: larger runners, pipeline complexity, and ongoing OOM risk. Managed platforms like Scalr run Terraform in optimized environments with a shared provider cache, and self-hosted agents let you set the run container's memory limit, so less of the tuning work lands on your pipeline.

About the author

Sebastian Stadil

CEO at Scalr

Sebastian Stadil is the CEO of Scalr with 15+ years of DevOps experience. He started with AWS in 2004 and advised early Microsoft Azure and Google Cloud.

Part of this guide

9 sheets

Terraform Troubleshooting, Optimization and Error Resolution

8 articles