Home › Guides › Apache Airflow Tutorial
Tutorial · 2026Apache Airflow Tutorial 2026: How Airflow 3 Works, Your First DAG and What It Really Costs
Apache Airflow is an open-source orchestrator that runs data pipelines written as Python code. You declare tasks and their dependencies in a DAG, and Airflow decides when each one runs, retries failures and records every attempt. The current release is 3.3.2, published on 17 September 2026. Self-hosting costs nothing in licence fees.
- Airflow is a scheduler, not a processing engine. It decides when code runs and what happens when it fails. The heavy lifting belongs in SQL, Spark or your warehouse.
- Version 3.3.2 landed on 17 September 2026 and splits parsing, scheduling and execution into separate processes, so a slow DAG file no longer stalls the scheduler.
-
A working UI takes about 20 minutes, using either
airflow standaloneor the Astro CLI, with no cloud account involved. - Licence cost is zero, operating cost is not. Amazon MWAA starts near $0.49 an hour for a small environment, billed per second.
- Asset-driven scheduling is the one Airflow 3 feature worth rewriting old DAGs for: it replaces guessed cron offsets between pipelines with a real dependency.
- Do not install Airflow for one nightly script. Below about five interdependent steps, cron and a good alert win.
- dbt is the tool most often paired with Airflow, at 44% of respondents in the Apache Airflow Survey 2025. Learn them together.
Picture the two-person data team at a mid-size insurer in Pune. Their nightly claims extract ran fine for four months, then last Tuesday it silently produced half a table, and nobody noticed until Wednesday's board pack showed the wrong number. Cron did its job perfectly: it started a script at 01:30. What it could not do was say which of eleven steps died, retry just that step, or refuse to load a table that came back almost empty. That gap is why Airflow exists.
What Is Apache Airflow, and What Does It Actually Replace?
Airflow is a Python application that keeps a calendar and a memory. You write a file saying "these five things happen, in this order, at this time". Airflow turns that into scheduled runs, keeps the state of every attempt in a database, and gives you a web page to look at when something breaks. That file is a DAG, a directed acyclic graph, which is a formal way of saying "steps with arrows between them and no loops".
The phrase people reach for is "workflow orchestration", which explains nothing. The honest version: Airflow replaces the pile of cron entries, bash wrappers and Slack webhooks that every data team grows by accident in its second year.
| What you need | cron on a VM | Apache Airflow |
|---|---|---|
| Run a script at 01:30 every day | Yes, one line | Yes, schedule="30 1 * * *"
|
| Know which of eleven steps failed | No, you get one exit code | Per-task state and per-task logs in the UI |
| Retry only the step that broke | You write the retry logic yourself |
retries=2 on the task, nothing else |
| Stop the load when the extract returned nothing | No, the next command runs regardless | Raise an exception and downstream tasks never start |
| Rerun the last 90 days in order | A bash loop and a lot of hope |
airflow dags backfill with a date range |
| Show a manager why Tuesday was late | Read syslog together | Gantt and grid views, one URL |
| Start when a file lands, not at a fixed time | Not without writing a watcher | Asset-driven scheduling, built in since Airflow 3 |
Read that as a threshold test, not a sales pitch. If your answer to most rows is "cron is fine", cron is fine. The insurer's team crossed the line on row four: they needed the load step to refuse to run, and no amount of bash was going to make that safe.
How widely Airflow is actually used
The real reason to learn Airflow before a newer orchestrator: almost every error you hit has already been answered somewhere.
Figures published by the Apache Airflow project in its Airflow Survey 2025 write-up, checked 24 September 2026.
Orchestration is a DP-700 skill area, so it shows up early in 360DT's live Microsoft Fabric data engineering course, where pipelines get built and scheduled in class rather than described on a slide. If SQL and Power BI are still shaky, the data analyst track is the more honest starting point.
Also read: SQL Query Optimization in 2026: 10 Mistakes That Make Queries Slow and How to Fix Each One, because most slow pipelines are slow queries wearing an orchestrator costume.
How Apache Airflow Works: The Five Components That Actually Run
Most tutorials hand you a diagram with a webserver in it. That diagram is now wrong. Airflow 3 rearranged the internals in a way that matters the first time you deploy, so learn the current shape rather than unlearning the old one later.
The five running components of an Airflow 3 install
Notice that the worker never touches the database. That single change is most of what Airflow 3 is about.
Component roles follow the Airflow 3.3.2 architecture overview in the official documentation, checked 24 September 2026.
Five processes, one folder. Walk it in order. The DAG processor reads your Python files and writes a serialised version of each DAG into the metadata database. In Airflow 3 it is its own process, so a file that takes nine seconds to import no longer stalls scheduling. The scheduler reads those serialised DAGs, works out which runs are due, and queues task instances. A worker picks one up and runs your function.
Then the change worth knowing in an interview. In Airflow 2, workers connected straight to the metadata database. In Airflow 3 they do not: task code talks to a Task Execution API served by the API server, which is also the UI, merged into one FastAPI process with a React front end. So you can run workers where the database is not reachable, and a task can no longer run SQL against Airflow's own tables.
Here is what usually goes wrong, and nobody warns beginners. Everything at the top level of a DAG file, outside a task, runs every single time the file is parsed, which is every few seconds by default. Put a requests.get() or a pd.read_csv() up there and you have built a machine that hammers an API forever while looking idle. Keep the module body to imports, constants and the DAG definition.
Apache Airflow Tutorial: Get a Running UI in 20 Minutes
Two routes. The choice is genuinely just whether Docker already works on your machine.
Route 1: pip, a virtual environment and one command
The official install docs insist on a constraints file, and they are right: Airflow pins a large dependency tree, and skipping constraints is how you get a broken environment ten minutes later.
python3 -m venv airflow-env
source airflow-env/bin/activate
export AIRFLOW_VERSION=3.3.2
export PYTHON_VERSION=3.12
pip install "apache-airflow==${AIRFLOW_VERSION}" \
--constraint "https://raw.githubusercontent.com/apache/airflow/constraints-${AIRFLOW_VERSION}/constraints-${PYTHON_VERSION}.txt"
export AIRFLOW_HOME=~/airflow
airflow standalone
airflow standalone initialises the database, creates a user and starts the components for you. It prints the generated password in the terminal, so do not close it. Open localhost:8080 and you have Airflow. This default uses SQLite, which means no parallel task execution: fine for learning, never for anything real.
Route 2: the Astro CLI, if you have Docker
This is the route I would pick in 2026: the same component layout production has, in throwaway containers.
mkdir airflow-claims && cd airflow-claims
astro dev init # writes dags/, Dockerfile, requirements.txt
astro dev start # builds the image, starts every component
astro dev logs -s # follow the scheduler log when something stalls
astro dev stop
astro dev start builds your project into an image, starts a container for each Airflow component plus a Postgres metadata database, and serves the UI on port 8080. There is a standalone mode too, which runs the components in a virtual environment instead. Containers, ports and volumes are the real prerequisite here, so our Docker tutorial is a sensible detour if docker ps is unfamiliar.
Also read: dbt Tutorial for Beginners 2026: Build and Test Your First dbt Model in 8 Steps. Airflow schedules, dbt transforms, and the survey data says most teams run them together.
Your First DAG: A Worked Apache Airflow Example You Can Run Tonight
This is the insurer's pipeline, reduced to three tasks that pass a value down the chain. Save it as dags/claims_daily.py. Note the import path: in Airflow 3 the public authoring interface is airflow.sdk, not the internal modules older blog posts still show.
# dags/claims_daily.py
from datetime import datetime
from airflow.sdk import dag, task
@dag(
dag_id="claims_daily",
schedule="30 1 * * *", # 01:30, after the source system closes
start_date=datetime(2026, 9, 1),
catchup=False,
default_args={"retries": 2},
tags=["claims", "finance"],
)
def claims_daily():
@task
def extract() -> int:
# replace with a real query; the return value is stored as XCom
rows = 4821
return rows
@task
def validate(rows: int) -> int:
if rows < 1000:
raise ValueError(f"only {rows} claims rows, the source is late")
return rows
@task
def load(rows: int) -> None:
print(f"loaded {rows} rows into warehouse.claims_fact")
load(validate(extract()))
claims_daily()
Four details are doing real work. catchup=False stops Airflow backfilling every day since 1 September, the single most common beginner surprise. retries: 2 gives a flaky network call two more chances without a human. The return value of extract reaches validate through XCom automatically, which is what the TaskFlow style buys you. And validate raises rather than returns, so load never runs on a short extract. That is the behaviour cron could not give the insurer.
Run it without waiting for 01:30
airflow dags list
airflow dags test claims_daily 2026-09-24
Here is the worked end to end. Input: 4,821 rows from extract. Steps: validate compares 4,821 against the 1,000-row floor and passes it through; load prints the row count. Output: three green squares in the grid view and a log line reading loaded 4821 rows into warehouse.claims_fact. Now change rows to 412 and run it again. You get a failed validate, a ValueError saying the source is late, two automatic retries, and a load task that never starts. That failure is the product. Everything else is plumbing.
Asset-Driven Scheduling, and Why It Is Worth Rewriting DAGs For
Most Airflow estates contain a lie that looks like this: DAG A runs at 02:00, DAG B runs at 02:30, and 02:30 was picked because A "usually" takes twenty minutes. The day A takes forty, B reads stale data and nobody is alerted, because as far as Airflow knows both DAGs succeeded.
Asset-driven scheduling deletes the guess. You declare that a DAG produces an asset, and the consumer is scheduled on that asset instead of the clock.
from airflow.sdk import asset
@asset(schedule="@daily")
def claims_raw():
# lands the raw extract; Airflow records that this asset was updated
...
@asset(schedule=claims_raw)
def claims_curated(context):
# runs because claims_raw was updated, not because the clock said 02:00
...
The 3.3.2 documentation describes @asset as shorthand for a DAG with one task that emits events for a single asset, and you can schedule on several at once with schedule=[Asset("a"), Asset("b")], combined using & and |. Airflow 3.3 also added a Task and Asset State Store, so tasks keep durable state across retries, plus a Language Task SDK that lets task logic be written in Java or Go while orchestration stays Python.
My call, and the trade-off I accept: convert your cross-DAG cron offsets to assets and leave everything else alone. The cost is a rewrite plus a week of the team relearning where a schedule is defined. Worth it, because those offsets are where silent staleness lives. What assets will not do is rescue a badly cut pipeline. If one DAG does nine unrelated things, assets just give you a prettier dependency on a mess.
The pattern travels. A retrieval pipeline that chunks and embeds documents is a batch job with a freshness contract, which is exactly an asset, and it is why orchestration sits inside the AI Engineer course on RAG and agents. Scheduled retraining and drift checks are the AI-300 territory covered by the MLOps engineer course.
What Apache Airflow Really Costs, From One VM to Managed
"Free and open source" is true about the licence and misleading about the bill. Airflow is five or six processes plus a Postgres database that must not be lost, and someone owns that.
Airflow at a glance, September 2026
Pin these four numbers before you argue with anyone about orchestration cost.
Release data from the Apache Airflow release notes; MWAA entry pricing from the AWS pricing page, both checked 24 September 2026.
| How you run it | What you pay | What you still own | Who it suits |
|---|---|---|---|
| Self-hosted, one VM, SQLite | Rs 0 licence plus the VM | Everything, and SQLite means no parallelism | Learning only, never production |
| Self-hosted, VM plus managed Postgres | Rs 0 licence plus VM and database | Upgrades, backups, TLS, the 2am page | One small team with an ops-minded engineer |
| Amazon MWAA | From about $0.49 an hour for a small environment, billed per second, plus storage | DAG code, IAM roles, VPC design | Teams already standardised on AWS |
| Airflow on Kubernetes, self-managed | Rs 0 licence plus your cluster | Helm values, autoscaling, node pressure | Platform teams who already run Kubernetes well |
| Commercial managed Airflow | Quoted per deployment, published on each vendor pricing page | DAG code, and almost nothing else | Teams who want Airflow without Airflow ops |
The number to hold anyone to is the managed entry point. AWS lists Amazon MWAA from about $0.49 an hour, billed per second, with environment classes from mw1.micro to mw1.2xlarge, and as of September 2026 MWAA supports Apache Airflow 3.3.1. At the entry size that is a few thousand rupees a month before storage, which is cheaper than most people guess and far cheaper than an engineer's weekend.
My recommendation is unromantic: self-host on one VM while you are learning, and move to managed the moment the pipeline matters to somebody other than you. Running your own Airflow teaches you a lot about Airflow and almost nothing about your data, which is a bad trade once a board pack sits downstream. If the cloud platform is the real gap, the AWS Solutions Architect and DevOps track and the Azure AZ-305 and AZ-400 track are where VPCs and IAM stop being mysterious, and the full certifications overview shows how they stack.
Learn orchestration the way a data engineering job actually uses it
The Microsoft Fabric Data Engineer Course covers Fabric data engineering for the DP-700 certification with DP-900 fundamentals included, over 8 live weekends. Hands-on projects, mentor support and placement guidance are part of the programme, and the live product page lists a batch starting 27 Sept 2026.
Explore the course
Where Airflow Breaks: 5 Failures You Will Hit in Month One
None of these are exotic. All five will happen to you, and each has a one-line fix once you recognise the symptom.
| What you see | What is actually happening | Fix |
|---|---|---|
| The DAG does not appear in the UI at all | An import error in the file, or the file is outside the DAG bundle path | Run airflow dags list-import-errors, then fix the traceback it prints |
| Hundreds of runs fire the moment you unpause a new DAG |
catchup defaults to filling every interval since start_date
|
Set catchup=False unless you genuinely want the backfill |
| The whole install is slow and the UI lags | Top-level code in a DAG file, such as an API call or a Pandas read, runs on every parse | Move that work inside a task; keep the module body to imports and definitions |
Tasks sit in queued and never start |
No worker capacity, or concurrency limits reached | Check worker logs first, then max_active_tasks and pool size |
| The worker is killed with no Python traceback | Out of memory, usually a DataFrame the size of the table | Push the transform into SQL or Spark and let the task only orchestrate it |
Airflow is a bad choice more often than its popularity suggests. It cannot do streaming, and pretending otherwise with a five-minute schedule is how teams end up with 288 half-broken runs a day. It is heavy to develop against locally, needing Docker and several gigabytes. Its UI is for engineers, not for business users who want to trigger a refresh. And if your pipeline is genuinely one script that runs at night, installing Airflow means you now operate a distributed system to solve a problem cron already solved. Use it when you have dependencies between steps, a real need for retries, and somebody who will look at the UI.
For the far end of the curve: DoorDash runs one of the largest Airflow deployments anywhere and, per Astronomer's published case study, built an internal layer called Orchestration Frederator so the platform scales horizontally by adding instances behind it, with no change required from DAG owners. The lesson for a two-person team is not to copy that architecture. It is that Airflow's limits at scale are DAG parsing, scheduler memory and metadata database connections, which is exactly what the Airflow 3 component split attacks.
How to Learn Apache Airflow in Four Weekends
Airflow rewards building one pipeline properly far more than watching ten hours of video. The plan below assumes about six hours a weekend, roughly what a working analyst in Bengaluru with a demanding job can actually find.
Four weekends to a DAG you would show an interviewer
Roughly six hours per weekend. Each stage ends in something you can open, run or link to.
Install, break, reinstall
Run airflow standalone, open port 8080, trigger an example DAG and read its logs. Deliverable: a green run plus a note on what the scheduler log said while it waited.
Write three tasks that pass data
Build this guide's extract, validate, load DAG against a real CSV or a free API. Deliverable: a DAG that fails loudly when the row count drops, proven with airflow dags test.
Add a warehouse and a connection
Point the load task at Postgres using an Airflow connection instead of a hardcoded string, then schedule it daily with catchup=False. Deliverable: seven days of run history in the UI.
Convert one schedule to an asset
Replace a guessed cron offset between two DAGs with asset-driven scheduling, then delete the offset. Deliverable: a README saying why the second DAG no longer needs to know when the first finishes.
Put it somewhere other people can see
Push the repo with a diagram, sample logs and the failure you caused on purpose. Deliverable: one GitHub link to paste into an application instead of writing "familiar with Airflow" on a CV.
Time estimates are our own teaching experience, not a vendor claim. Checked 24 September 2026.
Two things people skip and should not. Read one scheduler log line by line, because that is where every "why is nothing running" question is answered. And break a task deliberately so you can watch a retry happen: one engineered failure teaches more than six green runs. If seeing it done once helps, 360DT runs free webinars where pipelines get built live and you can ask why a choice was made.
What I Would Do in Your Position
If you are an analyst or developer moving into data engineering this year: install Airflow this weekend, build the claims DAG, break it on purpose, and stop there. Skip Kubernetes deployments, custom operators and home-grown plugins. Four weekends of one honest pipeline, with a README explaining why catchup=False is set and why the validate task raises, beats a certificate listing twelve tools.
Then decide what you actually want the job to be. Airflow is the portable, employer-agnostic skill; Microsoft Fabric is the one showing up in Indian job posts with a certification attached to it. If that second path is yours, the Microsoft Fabric Data Engineer course is the direct route, and a free demo class is a cheaper way to find out than a refund request.
Related guides
- Data Engineer Roadmap 2026: 7 Steps to Land Your First Job in India With Microsoft Fabric once you can build a DAG, this is the order to learn everything around it.
- Microsoft Fabric vs Databricks in 2026: Cost, Certifications and Which One Gets You Hired in India the platform decision that determines what your DAGs will actually call.
- What Is Apache Iceberg in 2026? The Open Table Format Explained, Iceberg v3, and 7 Skills to Learn the table format your load task is increasingly writing into.
- Data Engineer Salary in India 2026: What 24,000+ Listings Pay by Experience and Company Type what the skill is worth before you spend four weekends on it.
- Data Engineer Jobs in Hyderabad 2026: What 4,000+ Openings Say About Salary and Skills a concrete market to check your stack against.
- Data Engineering Projects for 2026: 6 Portfolio Builds With Stack, Dataset and Real Costs six builds your new DAG can become the backbone of.
Frequently asked questions
Is this Apache Airflow tutorial enough to get started with no data engineering experience?
It is enough to get a running install and a working three-task DAG, which is the hard part of week one. You need comfortable Python, plus enough SQL to write the query your extract task will run. If either is shaky, fix that first; Airflow will not teach it to you.
Is Apache Airflow free to use?
The software is free under the Apache 2.0 licence, at any scale. What costs money is running it: a virtual machine, a Postgres metadata database, and your own time on upgrades and backups. Managed options remove that work for a fee, and Amazon MWAA lists an entry price of roughly $0.49 an hour for a small environment.
Do I need to know Python to learn Airflow?
Yes, and more than beginner Python. DAGs are Python modules, the Task SDK is decorator-based, and most debugging is reading a traceback in a task log. You do not need advanced Python: functions, imports, exceptions, f-strings and a rough sense of what a decorator does will carry your first six months.
What is the difference between a DAG and a task in Airflow?
A task is one unit of work, such as running a query or copying a file. A DAG holds tasks, the order they run in, and the schedule. One DAG run creates one task instance per task, and each has its own state, logs and retry count.
How long does it take to learn Apache Airflow?
Four focused weekends gets you to a DAG you would show an interviewer: install, a three-task pipeline, a real connection, one asset-driven schedule. Production competence, meaning you can size workers, tune parsing and diagnose a stuck queue, is more like four to six months on something that matters.
Can Airflow handle real-time or streaming data?
Not as a stream processor, and you should not try. Airflow schedules and monitors batch work; latency is minutes, not milliseconds. For real streaming use Kafka, Flink or a managed service, and let Airflow orchestrate the batch jobs around them: compaction, backfills, daily aggregates.
Is Airflow still worth learning in 2026, or has it been replaced?
It is still the default. The Apache Airflow Survey 2025 drew 5,818 responses from 122 countries, and the project reports over 3,600 unique contributors, more than Spark or Kafka. Newer orchestrators are genuinely nicer in places, but Airflow is the one on the job description.
Should I learn Airflow or Microsoft Fabric pipelines first?
Learn whichever your target employers run, and in India that often means both. Fabric pipelines show results faster and map directly to DP-700; Airflow is more portable and appears in more senior posts. With no constraint, start with Airflow because the concepts transfer, then add Fabric.
About this guide. 360 Digital Transformation is an Authorized Training Partner of Anthropic and Microsoft. Other certification bodies, vendors and employers named here are not affiliated with us. Tools and versions change quickly; commands and figures cited were checked on 24 September 2026.




