DevOps Engineering Bootcamp
“Build. Containerize. Automate.”
Go from “what is a terminal?” to a live, containerized, multi-service app on Kubernetes, backed by Terraform-managed infrastructure.
Coming Soon Enrollment opens soon. Join the waitlist to hear first.
Stackable, self-paced boot camps with direct access to T2S instructors. They take you from your first terminal command to a self-healing reliability platform on AWS. Each boot camp gets you job-ready for its role on its own, and each one prepares you for the next.
Prefer one-on-one? 1-on-1 Coaching gives you one live session a month on any boot camp topic.
1module "eks" {2 source = "terraform-aws-modules/eks/aws"3 version = "~> 20.0"4 cluster_name = "t2s-platform"5 cluster_version = "1.30"6 vpc_id = module.vpc.vpc_id7 subnet_ids = module.vpc.private_subnets8 9 eks_managed_node_groups = {10 default = {11 instance_types = ["t3.large"]12 min_size = 213 max_size = 514 }15 }16}
Engineering teams hire in a familiar order. First they need someone who can ship a service. Then someone who can keep it alive. Then someone who can automate the response so problems get caught before a human sees them. The series follows that same order.
“Build. Containerize. Automate.”
Go from “what is a terminal?” to a live, containerized, multi-service app on Kubernetes, backed by Terraform-managed infrastructure.
“Keep It Running. Prove It Recovers.”
Observe the platform, break it on purpose, detect the failure, respond correctly, and prove with evidence that it recovered.
“Automate the Fix. Not Just the Build.”
Teach the platform to detect, score, and fix its own failures, then ship one finished, interview-ready system on AWS.
Short foundations courses for a strong start, a parallel AI/ML engineering track, and an advanced track in building autonomous agents.
“Before Docker, Before Kubernetes, the Shell.”
Build real comfort at the Linux command line, so the DevOps Bootcamp can move at full speed.
“The Language Before the Model.”
Not general-purpose Python. This is the specific subset you use to move data, train a model, and call an LLM.
“From Zero to a Model in Production.”
Data pipelines, model training, model serving, and monitoring live models, taught the way production teams work. Prefer one-on-one? Request 1-on-1 coaching.
“Systems That Decide, Not Just Systems That Alert.”
Build a Super Intelligence (SI, formerly known as AI) agent that watches the platform, reasons about incidents, and fixes them, with guardrails suited to regulated industries.
Start with Cloud Engineering (Linux → DevOps → SRE) or AI/ML Engineering (Python → Zero to AI/ML Systems Engineering). Both tracks lead to AIOps, where DevOps, SRE, and SI/ML come together. The Agent Build Track is the capstone above it all.
Week-by-week modules, the capstone you'll ship, the tools you'll use, and the roles you'll be ready for.
“Before Docker, Before Kubernetes, the Shell.”
Every DevOps and SRE skill in this series assumes you're comfortable at a Linux command line. This course builds that comfort first, so Course 1 can move at full speed instead of stopping to explain chmod.
| Week | Module | What You Learn |
|---|---|---|
| 01 | Shell Fundamentals | Navigation, file permissions, users and groups, package managers, process management (ps, top, kill), systemd services |
| 02 | Networking & Scripting | Ports, DNS basics, curl/netcat, bash scripting, log inspection (journalctl, tail, grep), SSH and remote access |
Diagnose and fix a deliberately broken Linux service using only the command line. Then write two paragraphs on what failed, why, and what fixed it. Course 1 builds on that habit from day one.
$ systemctl status checkout.service● checkout.service - Checkout API Active: failed (Result: exit-code)$ journalctl -u checkout.service -n 1checkout[812]: Error: listen EADDRINUSE: address already in use :8080$ sudo systemctl restart checkout.service✓ checkout.service is active (running)
“The Language Before the Model.”
This isn't general-purpose Python. It's the specific subset you use to move data, train a model, and call an LLM. The course exists so Zero to AI/ML Systems Engineering can start on systems engineering, not syntax.
| Week | Module | What You Learn |
|---|---|---|
| 01 | Python for Data | Syntax, data structures, functions, virtual environments and pip, NumPy, pandas |
| 02 | Python for ML | scikit-learn basics, tensors in PyTorch/TensorFlow, Jupyter notebooks, calling LLM APIs (Anthropic/OpenAI SDKs) |
Train a simple model end to end, from data in to prediction out. Then call an LLM API to summarize the result in plain language.
1import pandas as pd2from sklearn.model_selection import train_test_split3 4df = pd.read_csv("transactions.csv")5X = df.drop(columns=["is_fraud"])6y = df["is_fraud"]7 8X_train, X_test, y_train, y_test = train_test_split(9 X, y, test_size=0.2, random_state=4210)11print(len(X_train), "training rows")
“Build. Containerize. Automate.”
Go from “what is a terminal?” to deploying a live, containerized, multi-service application on Kubernetes, backed by Terraform-managed infrastructure. You build each version by hand first, then automate it, the same way real engineering teams work.
| Week | Module | What You Build |
|---|---|---|
| 01 | Local Foundations | A Node/Express service running on your own laptop |
| 02 | Containerization | The app split into three Docker services (Flask, Node, Web UI) |
| 03 | Manual Cloud Deployment | Docker Compose orchestration, then a hand-built AWS deploy (ECR, ECS/Fargate, ALB, VPC) |
| 04 | Observability Basics & IaC | Prometheus/Grafana dashboards; the ECS stack rebuilt from Terraform |
| 05 | Kubernetes Fundamentals | Workloads on Amazon EKS: probes, autoscaling, self-healing |
| 06 | Infrastructure as Code at Scale | Reusable Terraform modules, dev/prod environments, tagging, cost budgets |
Deploy a working three-service application to Kubernetes on at least one cloud, backed by modular, reusable Terraform. Then tear it down cleanly on command.
$ terraform apply -auto-approvemodule.eks.aws_eks_cluster.this[0]: Creation complete after 9m12sApply complete! Resources: 42 added, 0 changed, 0 destroyed.$ kubectl get pods -n checkoutNAME READY STATUS RESTARTScheckout-api-7c9f6d-2xkqp 1/1 Running 0checkout-api-7c9f6d-9hzt4 1/1 Running 0
“Keep It Running. Prove It Recovers.”
DevOps gets the system live. SRE keeps it alive. Bring the platform you built in Course 1, or one from your own experience. You'll learn to observe it, break it on purpose, detect the failure, respond correctly, and prove with evidence that it recovered.
| Week | Module | What You Build |
|---|---|---|
| 01 | Incident Management Foundations | Runbooks, alert routing, and incident response scoped to blast radius |
| 02 | Reliability Governance & Policy Gates | GitOps with ArgoCD; security and policy scanning with Trivy, OPA Gatekeeper, and Checkov to block unsafe deploys before they reach the cluster |
| 03 | Chaos Engineering & Incident Lifecycle | Controlled failure injection, Slack/ServiceNow/Jira ticket automation, blameless postmortems |
| 04 | Deep Observability & Error Budgets | SLOs, error budgets, on-call rotation practices, and a production-depth review of Prometheus, Grafana, and OpenTelemetry |
| 05 | Incident Response Drill Week | A full chaos drill you run yourself, from first alert to finished postmortem |
Take a running service, inject a controlled failure, and run the full incident lifecycle end to end: alert, ticket, runbook, recovery, and a documented postmortem with MTTR evidence.
1- alert: HighErrorRate2 expr: |3 sum(rate(http_requests_total{status=~"5.."}[5m]))4 / sum(rate(http_requests_total[5m])) > 0.025 for: 5m6 labels:7 severity: page8 annotations:9 runbook: "runbooks/checkout-api.md"
“From Zero to a Model in Production.”
This is the SI/ML counterpart to the DevOps Bootcamp. It's not a notebook full of experiments. It's how production teams engineer and run SI/ML systems: data pipelines, model training, model serving, and monitoring models once they're live. The final module covers responsible, auditable SI for regulated industries, drawn from applied research on SI and ML in those settings.
| Week | Module | What You Build |
|---|---|---|
| 01 | ML Systems Foundations | The ML lifecycle end to end: data in, model out, evaluated and reproducible |
| 02 | Data Pipelines & Feature Engineering | Ingesting, cleaning, and versioning data; a basic feature store |
| 03 | Model Training & Evaluation at Scale | Training pipelines, experiment tracking, hyperparameter tuning |
| 04 | Model Serving & Inference Infrastructure | Containerizing a model behind an API, GPU vs. CPU tradeoffs, a vector database for retrieval (RAG) |
| 05 | MLOps | CI/CD for models, a model registry, drift detection, retraining triggers |
| 06 | Responsible & Regulated SI | Governance, auditability, and safety patterns for SI running in fintech and healthcare systems |
Deploy a trained model as a live inference service with monitoring, a model registry entry, and a drift-detection alert. It's the MLOps counterpart to the DevOps capstone.
1import mlflow2from sklearn.ensemble import RandomForestClassifier3from sklearn.metrics import f1_score4 5mlflow.set_experiment("fraud-detection")6 7with mlflow.start_run():8 model = RandomForestClassifier(n_estimators=200)9 model.fit(X_train, y_train)10 score = f1_score(y_test, model.predict(X_test))11 mlflow.log_metric("f1", score)12 mlflow.sklearn.log_model(model, "model")
“Automate the Fix. Not Just the Build.”
This is where DevOps and SRE come together. You stop fixing incidents by hand and teach the platform to detect, score, and fix its own failures. Then you combine everything from Courses 1 and 2 into one finished, interview-ready system on AWS.
| Week | Module | What You Build |
|---|---|---|
| 01 | SI-Powered Detection | Anomaly detection, automated risk scoring, and SI-assisted incident summaries added to the Course 2 alert pipeline |
| 02 | Self-Healing Automation | Automated recovery scripts for common failure modes, a recovery policy loop, and a chaos suite that proves it worked |
| 03 | AWS Capstone Build | One standalone platform with the application services, CI/CD, GitOps, AIOps, observability and alerting, and FinOps, deployed and documented on AWS |
| 04 | Career Capstone | Portfolio review, GitHub profile, LinkedIn, resume, and interview prep, all built around your finished platform |
A fully self-healing reliability platform on AWS: the complete Express Reliability Platform. You present it as an interview-ready portfolio piece and walk a T2S instructor through it for review.
1def triage(alert):2 signals = correlate(alert, window_minutes=15)3 runbook = find_runbook(alert.service)4 risk = score_risk(signals)5 6 # low-risk, known fix: act; otherwise hand a human the summary7 if risk < 0.3 and runbook.auto_fix:8 return apply_fix(runbook, alert)9 return page_on_call(alert, summarize(signals))
“Systems That Decide, Not Just Systems That Alert.”
Every earlier course teaches you to detect and fix problems faster. This track teaches you to build the agent that does it for you. The agent plugs into the same platform you built across the series. It watches for an incident, reasons about the right response, and then either carries it out or hands a human a ready-made recommendation, with guardrails suited to regulated industries.
| Week | Module | What You Build |
|---|---|---|
| 01 | Agent Architecture Foundations | What makes a system an “agent”: perception, planning, tool use, and memory; LLM tool-calling basics |
| 02 | Building the Tool Layer | Safely connecting an agent to real systems from the series (kubectl, Terraform, Slack, ServiceNow) with permissions and sandboxing |
| 03 | Guardrails for Regulated Environments | Human-in-the-loop approval gates, audit logging, safe rollback, and governance patterns for agents running in fintech and healthcare systems |
| 04 | Capstone Agent Build | A working agent that spots a platform incident and proposes a fix, or carries out an approved one |
A live demo: your agent detects a simulated incident on the platform and reasons about the right fix. Then it either carries out the fix under guardrails or hands the on-call engineer a ready-made recommendation. You present it to a T2S instructor for review.
1tools = [{2 "name": "restart_deployment",3 "description": "Restart a Kubernetes deployment",4 "input_schema": {5 "type": "object",6 "properties": {7 "namespace": {"type": "string"},8 "deployment": {"type": "string"},9 },10 "required": ["namespace", "deployment"],11 },12}]
Each month you get one live, one-on-one session with a senior T2S instructor on a topic from the T2S boot camps. You choose the topic, we work through it together, and you leave with a clear plan for the month ahead.
1- alert: HighErrorRate2 expr: |3 sum(rate(http_requests_total{status=~"5.."}[5m]))4 / sum(rate(http_requests_total[5m])) > 0.025 for: 5m6 labels:7 severity: page8 annotations:9 runbook: "runbooks/checkout-api.md"
Share your background, goals, and the topic you want to start with using this form.
We'll reach out to talk through fit, topics, and pricing.
Pick your topic, come with your questions, and leave with a plan for the month.
One live session a month on any boot camp topic. We'll reply by email, usually within two business days.
Every version, every course, and every cloud follows the same loop, from week 1 of DevOps through the last day of AIOps.
Read the purpose and key concepts before touching anything.
Follow exact commands in exact order.
Confirm expected output at every step.
Cause a failure on purpose in a safe, controlled environment.
Use real tools (logs, metrics, alerts) to restore service.
Write down what failed, why, and what fixed it.
Turn the fix into a script so you never do it by hand again.
Make the system harder to break and faster to recover.
Every course in the series is built on AWS. Instead of skimming several providers, you go deep on one, so you can speak with confidence about the AWS services employers ask about.
| Concept | What you use on AWS |
|---|---|
| Managed Kubernetes | Amazon EKS |
| Containers without Kubernetes | Amazon ECS on Fargate |
| Container Registry | Amazon ECR |
| Networking & Load Balancing | VPC and Application Load Balancer |
| IaC Provider | hashicorp/aws (Terraform) |
| CLI Tool | aws |
You never have to commit to the full pathway up front. Each course stands on its own, and each one gives you exactly what the next course needs, so you never feel like you're starting over.
| After Completing | You Can Credibly Apply For | You Hold |
|---|---|---|
| Course 1 only | DevOps Engineer, Cloud Engineer, Platform Engineer (entry) | DevOps Engineering Bootcamp Certificate |
| Courses 1 + 2 | + Site Reliability Engineer, Incident Response Engineer, DevSecOps Engineer | + SRE Bootcamp Certificate |
| Courses 1 + 2 + 3 | + AIOps Engineer, Senior SRE, Platform/Reliability Architect, AWS Cloud Architect | T2S Cloud Reliability Engineer Pathway Credential (all three certificates) |
| Zero to AI/ML Systems Engineering (with Python for AI and ML) | ML Engineer, AI/ML Systems Engineer, MLOps Engineer, Applied AI Engineer | Zero to AI/ML Systems Engineering Certificate |
| Agent Build Track (after Courses 1–3, the AI/ML track, or equivalent) | + AI Agent Engineer, Applied AI/Automation Engineer | + Agent Build Track Certificate, the highest tier of the Pathway Credential |
Pay once and save, or spread the cost over a short plan with no interest. Every course and track includes:
| Course / Track | Suggested Pace | Single Payment | Installment Plan |
|---|---|---|---|
| Linux for DevOps and SRE (Foundations) | 2 weeks | $249 · save $21 | 3 × $90/mo ($270) |
| Python for AI and ML (Foundations) | 2 weeks | $249 · save $21 | 3 × $90/mo ($270) |
| Course 1: DevOps Engineering Bootcamp | 6 weeks | $1,795 · save $100 | 5 × $379/mo ($1,895) |
| Course 2: Site Reliability Engineering Bootcamp | 5 weeks | $1,995 · save $100 | 5 × $419/mo ($2,095) |
| Zero to AI/ML Systems Engineering | 6 weeks | $1,795 · save $100 | 5 × $379/mo ($1,895) |
| Course 3: AIOps Bootcamp | 4 weeks | $2,250 · save $100 | 5 × $470/mo ($2,350) |
| Agent Build Track (Signature Advanced) | 3–4 weeks | $2,450 · save $100 | 5 × $510/mo ($2,550) |
Linux Foundations + DevOps Bootcamp + SRE Bootcamp
$4,039 separately · Save $744
Bundle pricing goes to the waitlist first.
Join the Waitlist →Python Foundations + Zero to AI/ML Systems Engineering
$2,044 separately · Save $249
Bundle pricing goes to the waitlist first.
Join the Waitlist →All 7 courses and tracks: both foundations, DevOps, SRE, AI/ML Systems, AIOps, and the Agent Build Track
$10,783 separately · Save ~$3,788
or 6 × $1,225/mo ($7,350 total)
Bundle pricing goes to the waitlist first.
Join the Waitlist →Contact T2S for launch dates and any active promotions.
No. Each course stands on its own. You can take just the DevOps Engineering Bootcamp and leave ready to apply for entry-level DevOps roles. If you keep going, each course gives you exactly what the next one needs.
The DevOps Engineering Bootcamp has no prerequisites. It's built for people starting from zero. If you'd like a head start, take Linux for DevOps and SRE first so Course 1 can move at full speed.
If you want to build, deploy, and run infrastructure, start with Cloud Engineering (Linux → DevOps → SRE). If you want to build and run machine learning systems, start with AI/ML Engineering (Python → Zero to AI/ML Systems Engineering). Both tracks meet at the AIOps Bootcamp.
Yes. The SRE Bootcamp accepts equivalent working knowledge of containers, Kubernetes, and Terraform in place of Course 1. The AIOps Bootcamp and Agent Build Track accept working engineers with equivalent experience. Tell us about your background and we'll help you pick the right starting point.
No. Every course and track is self-paced, so you can fit it around your job and your life. You're never on your own: you get direct access to T2S instructors. If you want live time too, add 1-on-1 Coaching.
Yes. 1-on-1 Coaching gives you one live session a month with a senior T2S instructor on a topic you choose from any boot camp, plus email support between sessions. Fill out the coaching form and we'll reach out to talk through fit, topics, and pricing.
AWS. Every course is built on it, from ECR and ECS on Fargate to EKS and Terraform's AWS provider, so you can speak to it with confidence in an interview.
At T2S, we call the field Super Intelligence (SI), formerly known as Artificial Intelligence (AI). Course names, job titles, and industry terms such as AIOps still say “AI,” so they match what employers post and what you'll search for.
Every course and track comes with a 14-day money-back guarantee. There's never an income-share agreement.
Enrollment opens soon. Because every course is self-paced, you can start as soon as it does. Join the waitlist and you'll get launch dates and pricing before anyone else.
Join the waitlist and we'll send you launch dates, pricing, and early-access details before they go public.