REAL PROBLEMS. REAL SOLUTIONS.
Not just wins — the real engineering. The bugs, the fires, the failures that made me better.
Kubernetes Pods in CrashLoopBackOff
PROBLEM
After a deployment, pods immediately entered CrashLoopBackOff with no obvious error in application logs. The service was unreachable.
INVESTIGATION
Ran `kubectl describe pod` — found OOMKilled in events. Checked resource limits — the container was requesting 128Mi but the app needed 512Mi under load. The previous deployment had worked only because traffic was low at the time.
SOLUTION
Updated the Deployment manifest with correct resource requests and limits (512Mi memory, 250m CPU). Added horizontal pod autoscaling and configured proper readiness/liveness probes to prevent traffic routing to unhealthy pods.
LESSON
Resource limits should reflect real measured usage, not optimistic estimates. Always test deployments under realistic load. Liveness probes are your last line of defense, not your first.
Secrets Exposed in CI/CD Pipeline Logs
PROBLEM
During a GitHub Actions pipeline run, AWS credentials were being printed to build logs — visible to anyone with repository access.
INVESTIGATION
Traced the issue to a debug `echo` statement left in a shell script that was printing environment variables. The variable happened to contain the AWS_SECRET_ACCESS_KEY passed as a GitHub Secret.
SOLUTION
Removed the debug statement immediately. Rotated the exposed AWS credentials in IAM. Audited all pipeline scripts for any env var printing. Added a step to scan for potential secret leakage using `git-secrets` and `truffleHog` in the pipeline itself. Implemented secret masking rules in GitHub Actions.
LESSON
Secrets rotation should be your first action when exposure is suspected, not your last. Debug statements in automation are a security risk. Scanning for secrets in CI is not optional for any production pipeline.
Docker Container Cannot Reach Database
PROBLEM
A Flask application container was failing to connect to a PostgreSQL container with 'Connection refused' — both containers appeared to be running.
INVESTIGATION
Checked Docker logs — PostgreSQL was starting but the Flask app was starting faster and attempting connection before Postgres was ready. Also discovered the Flask app was using `localhost` to connect instead of the Docker service name.
SOLUTION
Fixed the connection string to use the Docker Compose service name (`db`) instead of `localhost`. Added a health check on the PostgreSQL container and a `depends_on: condition: service_healthy` directive in Docker Compose so the app waits for the database to be ready.
LESSON
In Docker networking, containers communicate via service names, not localhost. Startup order and readiness are different problems — `depends_on` handles order, health checks handle readiness.