Book 05 · Cheat sheet
DevOps & Cloud
Tools and practices for building, testing, packaging and shipping software to production.
Software Dictionary · softwaredictionary.org/categories/devops/cheat-sheet
- 01Ansible
- Ansible is an open-source automation tool that configures servers and deploys applications by running YAML playbooks over SSH, with no agent on the machines.
- Ansible automates server configuration and application deployment.
- Tasks are written in YAML playbooks and run against an inventory of hosts.
- It is agentless: it connects over SSH and needs only Python on the hosts.
- 02Autoscaling
- Autoscaling is the automatic adding or removing of computing resources, such as servers or containers, based on demand to keep performance steady and costs low.
- Autoscaling adds or removes capacity automatically based on demand.
- It is driven by metrics such as CPU, memory, request rate, or queue length.
- Minimum and maximum limits keep scaling safe and costs predictable.
- 03AWSAmazon Web Services
- AWS (Amazon Web Services) is Amazon's cloud platform: more than 200 pay-as-you-go services for servers, storage, databases and much more.
- AWS is Amazon's cloud platform, with more than 200 pay-as-you-go services.
- Core services include EC2 (servers), S3 (storage), RDS (databases) and Lambda.
- Regions contain several availability zones for resilience.
- 04AzureMicrosoft Azure
- Microsoft Azure is Microsoft's cloud computing platform, offering hundreds of on-demand services such as virtual machines, databases and AI models worldwide.
- Azure is Microsoft's public cloud platform, launched in 2010.
- It is one of the three largest clouds, with AWS and Google Cloud.
- Services range from VMs and Kubernetes to databases, functions and AI models.
- 05Blue-Green Deployment
- Blue-green deployment is a release strategy that uses two identical production environments and moves all traffic from the old version to the new one at once.
- Two identical environments exist: one live, one idle.
- The new version is deployed and tested on the idle environment before going live.
- Traffic switches all at once, typically at the load balancer, with no downtime.
- 06Canary Deployment
- A canary deployment releases a new software version to a small share of users first, checks its health, and then gradually rolls it out to everyone.
- A canary sends a small percentage of traffic to the new version first.
- Traffic is increased in steps while metrics are watched.
- A bad release affects only a few users and can be rolled back quickly.
- 07CDNContent Delivery Network
- A CDN is a network of servers spread around the world that stores copies of website content and delivers it to each user from the nearest location.
- A CDN serves content from servers close to each user.
- It works by caching copies of files from the origin server.
- CDNs reduce latency, lower origin load, and help absorb traffic spikes and attacks.
- 08Chaos Engineering
- Chaos engineering is the practice of deliberately injecting failures into a system, such as crashing servers, to confirm that it keeps working as expected.
- Chaos engineering injects real failures on purpose to reveal hidden weaknesses.
- Each experiment starts from a measurable steady state and a clear hypothesis.
- Experiments keep the blast radius small and can be aborted at any time.
- 09CI/CDContinuous Integration / Continuous Delivery
- CI/CD is a set of automated practices that build, test, and release code changes frequently, so software can be delivered to users quickly and safely.
- CI automatically builds and tests every change merged into the shared codebase.
- CD keeps the code always ready to release, or releases it automatically.
- Pipelines are defined in configuration files stored alongside the code.
- 10Cloud Computing
- Cloud computing is the on-demand delivery of computing resources, such as servers, storage, and databases, over the internet with pay-as-you-go pricing.
- Cloud computing rents servers, storage, and services over the internet on demand.
- The three main service models are IaaS, PaaS, and SaaS.
- Elasticity lets resources grow and shrink automatically with demand.
- 11Configuration Management
- Configuration management is the practice of defining the desired state of servers and software in code and using tools to apply and keep it automatically.
- Configuration management keeps the desired state of systems in version-controlled code.
- Tools apply that state automatically and correct drift on existing machines.
- Idempotency makes it safe to apply the same configuration many times.
- 12Container
- A container is a lightweight, isolated package that bundles an application with its dependencies and runs it on the host's shared operating system kernel.
- A container packages an app with its dependencies so it behaves the same everywhere.
- Containers share the host operating system kernel, unlike virtual machines.
- Linux namespaces provide isolation, and cgroups limit resources.
- 13Container Registry
- A container registry is a storage and distribution service for container images, letting teams push built images and pull them onto any server that runs them.
- A registry stores container images and serves them to the machines that run them.
- Images are grouped into repositories and versioned with tags.
- A digest identifies an image's exact contents and cannot be moved like a tag.
- 14DevOpsDevelopment and Operations
- DevOps is a set of practices and a culture that brings software development and IT operations together to deliver software faster and more reliably.
- DevOps unites development and operations into one shared responsibility.
- Automation is central: CI/CD, infrastructure as code, and monitoring.
- Small, frequent releases reduce risk and speed up feedback.
- 15Distributed Tracing
- Distributed tracing is a technique that follows a single request as it travels through many services, recording how long each step took and where it failed.
- A trace follows one request end to end across services.
- Traces are made of spans, each measuring one operation, linked in parent-child order.
- A trace ID is passed between services, usually in the traceparent header.
- 16DNSDomain Name System
- DNS is the internet's naming system that translates human-readable domain names like example.com into the numeric IP addresses computers use to connect.
- DNS translates domain names into IP addresses.
- Resolvers find answers by querying root, top-level domain, and authoritative name servers.
- Common record types are A, AAAA, CNAME, MX, and TXT.
- 17Docker
- Docker is an open-source platform for packaging an application and everything it needs into a container that runs the same way on any machine.
- Docker packages an application and its dependencies into a container image.
- A Dockerfile defines how the image is built.
- An image is the template; a container is a running instance of it.
- 18Docker Compose
- Docker Compose is a tool for defining and running multi-container applications, such as a web server plus a database, from one YAML file with one command.
- Docker Compose runs multi-container applications from one YAML file.
- docker compose up starts every service on a shared network.
- Services declare images, ports, environment variables, volumes and health checks.
- 19Edge Computing
- Edge computing runs code and processes data close to where users or devices are, instead of in a distant central data center, to reduce latency.
- Edge computing runs code close to users or devices to reduce latency and bandwidth use.
- On the web, edge functions run in many locations, often on CDN networks.
- In IoT, edge devices process data locally and send only results to the cloud.
- 20Environment Variable
- An environment variable is a named value set outside a program, by the operating system or runtime, that the program reads to configure its behavior.
- Environment variables are key-value settings provided by the environment, not the code.
- They let the same code run with different configuration in each environment.
- Child processes inherit a copy of their parent's environment variables.
- 21Feature Flag
- A feature flag is a switch in code that turns a feature on or off at runtime, letting teams deploy code without releasing it to every user at once.
- A feature flag turns functionality on or off at runtime without a new deployment.
- It separates deploying code from releasing a feature to users.
- Targeting rules can enable a feature for specific users, groups, or a percentage of traffic.
- 22GitHub Actions
- GitHub Actions is GitHub's built-in automation platform: YAML workflows in a repo run tests, builds and deployments on events like a push or pull request.
- GitHub Actions runs automated workflows inside GitHub.
- Workflows are YAML files in .github/workflows, triggered by events.
- Jobs run on GitHub-hosted or self-hosted runners, step by step.
- 23GitOps
- GitOps is a way of managing infrastructure and deployments where Git holds the desired state of a system and an automated agent keeps the live system in sync.
- Git is the single source of truth for the desired state of the system.
- Changes go through commits and pull requests, not manual commands.
- An agent continuously reconciles the live system with the repository.
- 24Google CloudGoogle Cloud Platform
- Google Cloud is Google's public cloud platform, offering compute, storage, data and AI services on the global infrastructure behind Google's own products.
- Google Cloud is Google's public cloud platform, often called GCP.
- It started with App Engine in 2008.
- Compute Engine, Cloud Run and GKE run code; Cloud Storage and Cloud SQL store data.
- 25Grafana
- Grafana is an open-source tool for building dashboards that turn metrics, logs and traces from many data sources into live charts and alerts in one place.
- Grafana builds dashboards and alerts from existing data sources.
- It connects to Prometheus, Loki, Elasticsearch, SQL databases and more.
- Panels from different sources can be combined on one dashboard.
- 26Helm
- Helm is the package manager for Kubernetes: it bundles an app's configuration files into a chart you can install, upgrade and roll back with one command.
- Helm is the package manager for Kubernetes.
- A chart bundles templated manifests with default values in values.yaml.
- Each install is a release that can be upgraded and rolled back.
- 27IaaSInfrastructure as a Service
- IaaS (infrastructure as a service) is a cloud model where you rent virtual machines, storage and networks on demand and manage the OS and above yourself.
- IaaS rents virtual machines, storage and networks on demand.
- You manage the operating system, software, patches and scaling.
- Amazon EC2, launched in 2006, defined the model.
- 28Immutable Infrastructure
- Immutable infrastructure is an approach where servers are never changed after deployment; every update replaces them with new, freshly built ones.
- Servers and containers are never modified after they are deployed.
- Every change produces a new image, and old instances are replaced, not patched.
- The same tested image moves unchanged through staging and production.
- 29Infrastructure as Code
- Infrastructure as code is the practice of defining servers, networks, and other infrastructure in version-controlled files that tools apply automatically.
- IaC defines infrastructure in code files instead of manual console clicks.
- Declarative tools compare the desired state with reality and apply the difference.
- Infrastructure changes are versioned, reviewed, and rolled back like application code.
- 30Jenkins
- Jenkins is an open-source automation server that builds, tests and deploys software through pipelines, and one of the oldest and most widely used CI/CD tools.
- Jenkins is a self-hosted, open-source automation server for CI/CD.
- It forked from Hudson and took its name in 2011.
- Pipelines are defined in a Jenkinsfile with build, test and deploy stages.
- 31kubectl
- kubectl is the command-line tool for Kubernetes: it sends requests to a cluster's API server to deploy applications, inspect them and change them.
- kubectl is the Kubernetes command-line tool; every command is a call to the API server.
- get, describe, logs and exec inspect what is running.
- kubectl apply -f makes the cluster match YAML files that describe the desired state.
- 32Kubernetes
- Kubernetes is an open-source system that automates deploying, scaling, and managing containerized applications across a cluster of machines.
- Kubernetes orchestrates containers across a cluster of machines.
- You declare the desired state in YAML, and Kubernetes keeps it that way.
- Pods, deployments, and services are its core building blocks.
- 33Linux
- Linux is an open-source operating system kernel that powers most servers, cloud platforms, containers, and Android phones, usually packaged as a distribution.
- Linux is an open-source kernel; distributions add tools to make a full operating system.
- It runs most servers, cloud infrastructure, supercomputers, and Android devices.
- Containers depend on Linux kernel features such as namespaces and cgroups.
- 34Load Balancer
- A load balancer is a server or service that spreads incoming traffic across several backend servers so no single one is overloaded and the app stays available.
- A load balancer distributes requests across multiple servers.
- Health checks automatically take failing servers out of rotation.
- Common algorithms include round robin, least connections, and IP hash.
- 35Logging
- Logging is the practice of recording timestamped messages about events in a running program, such as errors and requests, so people can investigate them later.
- Logging records timestamped events from a running program for later investigation.
- Log levels such as DEBUG, INFO, WARN, and ERROR indicate severity and control volume.
- Structured logging writes entries as JSON with named fields, which makes them easy to query.
- 36Metrics
- Metrics are numeric measurements of a system collected over time, such as request rate, error rate and CPU usage, used for dashboards, alerts and planning.
- Metrics are numeric measurements recorded over time and stored as time series.
- The common types are counters, gauges, and histograms.
- Percentiles such as p95 and p99 reveal slow requests that averages hide.
- 37Object Storage
- Object storage is a way of storing data as whole objects, each with a unique key and metadata, in flat buckets that scale to huge numbers of files over HTTP.
- Object storage saves data as objects: bytes plus metadata, addressed by a unique key.
- Objects live in buckets with a flat namespace and are accessed over HTTP.
- Objects are replaced as a whole rather than edited in place.
- 38Observability
- Observability is the ability to understand what is happening inside a running software system by collecting and analyzing its logs, metrics, and traces.
- Observability means understanding a system's internal state from the data it emits.
- Logs, metrics, and traces are the three main types of telemetry.
- Distributed tracing follows one request across many services.
- 39OpenTelemetry
- OpenTelemetry is an open standard and set of tools for collecting traces, metrics and logs from software and sending them to any monitoring backend.
- OpenTelemetry is a vendor-neutral standard for traces, metrics and logs.
- Code is instrumented once, and the data can go to any backend.
- OTLP is its protocol; the Collector receives, processes and forwards data.
- 40PaaSPlatform as a Service
- PaaS (platform as a service) is a cloud model where you deploy your code and the provider runs everything under it: servers, operating systems and scaling.
- PaaS runs your code while the provider manages servers and the OS.
- Heroku popularized it; Vercel, Render and App Service are examples.
- Platforms add databases, scaling, logs, previews and rollbacks.
- 41Pod
- A pod is the smallest deployable unit in Kubernetes: one or more containers that share a network address and storage and are scheduled together on one node.
- A pod is the smallest unit that Kubernetes schedules and manages.
- Containers in the same pod share an IP address, localhost, and volumes.
- Most pods run one container; helper containers in the same pod are called sidecars.
- 42Postmortem
- A postmortem is a written review after an incident that explains what happened, why it happened, and what the team will change so it doesn't happen again.
- A postmortem analyzes an incident after it is resolved and records the lessons learned.
- It includes a timeline, impact, root causes, contributing factors, and action items.
- Blameless postmortems focus on system and process failures, not on individuals.
- 43Prometheus
- Prometheus is an open-source monitoring system that collects metrics from apps and servers, stores them as time series and alerts when values cross a limit.
- Prometheus is an open-source monitoring system and time-series database.
- It pulls metrics by scraping HTTP endpoints, usually /metrics.
- Each series has a name and labels; PromQL queries them.
- 44Reverse Proxy
- A reverse proxy is a server that sits in front of web servers, accepts client requests on their behalf, and forwards each request to the right backend server.
- A reverse proxy receives requests on behalf of backend servers and forwards them.
- It commonly handles HTTPS termination, caching, compression, and routing.
- Backend servers can stay hidden on a private network.
- 45Rollback
- A rollback is the process of returning software to a previous, known-good version after a new deployment causes errors, outages, or other unexpected problems.
- A rollback restores the last known-good version of an application after a bad release.
- It is usually the fastest way to reduce user impact during an incident.
- Blue-green deployments and versioned container images make rollbacks quick.
- 46SaaSSoftware as a Service
- SaaS (software as a service) is software delivered over the internet as a subscription; users sign in while the provider runs, updates and secures it.
- SaaS is software used over the internet as a subscription.
- The provider hosts, updates, secures and backs up the application.
- It is the top layer above PaaS and IaaS, with the least to manage.
- 47Serverless
- Serverless is a cloud model in which the provider runs your code on demand, manages all the servers, scales automatically, and bills only for actual use.
- Serverless does not mean no servers; it means you don't manage them.
- Code runs in response to events such as HTTP requests or file uploads.
- It scales automatically, including down to zero when idle.
- 48Service Mesh
- A service mesh is an infrastructure layer that manages traffic between microservices, adding encryption, retries, routing, and monitoring without code changes.
- A service mesh manages service-to-service traffic in a microservices system.
- Proxies form the data plane, and a control plane configures them.
- It provides mutual TLS, retries, timeouts, traffic splitting, and telemetry.
- 49Site Reliability Engineering
- Site reliability engineering is a discipline that applies software engineering to operations, keeping services reliable with automation and measurable targets.
- SRE applies software engineering to operations and infrastructure work.
- Reliability is measured with SLIs and SLOs that reflect what users experience.
- An error budget balances the pace of new releases against stability.
- 50SLAService Level Agreement
- An SLA (service level agreement) is a provider's commitment to customers about the level of service, such as 99.9% uptime, and what happens if it isn't met.
- An SLA is a formal promise about service quality, such as uptime.
- Missing it usually earns customers service credits.
- 99.9% allows about 43 minutes of downtime a month; 99.99% about 4.3 minutes.
- 51SLOService Level Objective
- An SLO is a measurable reliability target for a service, such as 99.9% of requests succeeding over 30 days, that tells a team how reliable is reliable enough.
- An SLO is a target value for a reliability measurement over a time window.
- The measurement itself is called an SLI, for example the share of successful requests.
- The allowed failure, 100% minus the SLO, is the error budget.
- 52Terraform
- Terraform is an infrastructure-as-code tool: you describe cloud resources in configuration files, and one command creates or updates them to match.
- Terraform describes cloud infrastructure as code, in HCL files.
- terraform plan previews changes; terraform apply makes them.
- A state file records which real resources Terraform manages.
- 53Virtual Machine
- A virtual machine is a software-based computer that runs its own operating system on shared physical hardware, isolated from other machines on the same host.
- A VM is a software-emulated computer with its own operating system.
- A hypervisor shares one physical host's hardware among many isolated VMs.
- A VM can run a different operating system than its host.
- 54YAMLYAML Ain't Markup Language
- YAML is a human-readable data format that uses indentation instead of brackets, widely used for configuration files in DevOps tools and CI/CD pipelines.
- YAML is a human-readable format for structured data and configuration.
- Indentation with spaces defines structure; tabs are not allowed.
- It supports comments with #, unlike JSON.