Cloud, DevOps & Infrastructure
Field notes on AWS at scale: multi-account guardrails, platform engineering with small teams, SecOps consolidation, and disaster recovery you have tested.
Field notes from running AWS for a regulated fintech: multi-account guardrails with a small platform team, cluster autoscaling, WAF in the agent era, consolidating SecOps telemetry, and ransomware recovery that starts with backups you've actually tested. The through-line is small-team leverage — platform engineering that lets a few people run a lot of infrastructure safely.
18 posts, newest first
Ransomware Recovery: A Tested-Backups Problem
Everyone has backups. Almost nobody has a restore they've actually run under fire. That gap is where ransomware turns a bad week into an existential one.
Guardrails at Scale for a Three-Person Team
A lean team can govern a sprawling cloud estate without becoming a ticket queue — but only if you put the rules in the pipeline, not in your inbox.
Internal IT as a Product, or Shadow IT Wins
Internal platforms fail when run like monopolies. Give them product managers, roadmaps, honest adoption metrics — and let users defect to better tools.
WAF in the Agent Era: Good Bots vs. Abuse
Agents are now real customers hitting your edge with real economics. The old bot question — human or machine? — is the wrong one. Here's the one that matters.
Consolidate SecOps on OCSF, Not Aggregators
Dashboard sprawl isn't a tooling gap you fix with more tooling — it's a schema problem. Standardize on OCSF and the single pane of glass becomes real.
Cluster Autoscaler to Karpenter: What Breaks
Karpenter is the right call for most EKS shops — but the migration breaks things unrelated to autoscaling. What to know before flipping the switch.
Autonomous Pentesting in a Regulated Shop
A tool that scans and exploits your estate on its own schedule is a gift and a loaded gun. The scoping, approvals, and evidence I'd want before it runs.
AWS Cost Levers That Moved the Needle
Cutting ~35% off a multi-region AWS footprint with no capability loss — the levers in the order they paid back, best first.
The Eight-Domain Azure Security Review
A tool scores your Azure posture; an assessor walks your architecture. The eight domains I review, in audit order, and the evidence each has to produce.
Zero-Downtime Database Changes Are a Process
Blue/green and serverless Aurora don't make migrations safe — the runbook does. The boring discipline that keeps schema changes from becoming incidents.
Aurora DSQL for the Ledger: Active-Active
Multi-region active-active sounds like the answer to ledger nightmares. Interrogate the consistency, recovery math, and migration before betting the books.
Fine-Grained Authorization for Fintech APIs
Authorization scattered across your codebase isn't a feature — it's a liability you can't prove. The pattern multi-tenant regulated platforms actually need.
MCP Is a New Attack Surface: An IAM Playbook
Every MCP server is a new identity reaching into your cloud. Whether that's leverage or liability comes down to least-privilege IAM on every tool call.
An SBOM Nobody Reads Is Compliance Cosplay
Generating a software bill of materials is the easy part. Wiring it into the moment a change ships is where supply-chain security stops being theater.
Warm Standby Is a Promise You Have to Test
A DR plan you have never exercised is a hypothesis with a logo on it. Warm standby only counts as a promise if you test the failover before you need it.
Security and DevOps Under One Roof
The case for running security and DevOps as one mandate: org-chart distance doesn't create security, and owning the pipelines changes how you protect them.
Post-Close Cyber Integration: A 100-Day Plan
The post-close decade is decided in the first 100 days. The eight cyber controls to ship by day 30, and the identity-sprawl audit every exit diligence will run.
Cloud FinOps: Where 25–40% of Spend Hides
The press-release version of cloud savings cancels workloads and books compliance debt. The durable version is commitment management and SaaS rationalization.