Data pipelines compared: when to use ETL, ELT or streaming, how to choose for each use case and the engineering practices that keep them reliable.
Every analytics or AI initiative depends on pipelines that move data reliably from source systems to where it is used. The three dominant patterns — ETL, ELT and streaming — suit different needs.
ETL: extract, transform, load
Data is transformed before it lands in the target system. ETL works well when the target has limited compute, when sensitive data must be filtered before loading, or when transformations are stable and well understood.
ELT: extract, load, transform
Raw data is loaded first and transformed inside a scalable warehouse or lakehouse. ELT keeps raw history, lets analysts change transformations without re-extracting data, and takes advantage of elastic cloud compute. It is now the default for many analytics teams.
Streaming
Events are processed continuously as they happen, often through a platform such as Apache Kafka. Streaming suits fraud detection, operational dashboards, IoT and any use case where minutes of delay matter. It adds complexity in ordering, state management and exactly-once processing.
Choosing an approach
| Need | Good fit |
|---|---|
| Daily or hourly reporting | ELT (or batch ETL) |
| Sensitive data that must be filtered first | ETL |
| Real-time decisions and alerts | Streaming |
| Change data from operational databases | Change data capture into ELT or streaming |
Engineering practices that matter
- Treat pipelines as code: version control, testing and code review.
- Add data quality checks and alert on failures and freshness.
- Document lineage so consumers know where data came from.
- Design for reprocessing — you will need to rerun history.
5 best practices for reliable data pipelines
- Define data contracts. Agree schemas, freshness and quality expectations with the teams that produce source data, so upstream changes do not silently break downstream reports.
- Make pipelines idempotent. Running a job twice should produce the same result. This makes retries and backfills safe.
- Monitor freshness and volume. Alert when data arrives late, when row counts change unexpectedly or when null rates spike, not just when a job fails.
- Separate orchestration from transformation. Use an orchestrator to schedule and track dependencies, and keep transformation logic in version-controlled, testable code.
- Document lineage. Automated lineage shows which dashboards and models depend on each source, which speeds impact analysis and incident response.
Choosing tools
Teams typically combine managed connectors for common SaaS sources, a transformation framework that runs SQL inside the warehouse or lakehouse, an orchestrator and, where needed, a streaming platform. Favour tools that integrate with your catalogue, support testing and fit your team’s skills.
Common mistakes to avoid
- Building streaming pipelines for use cases that only need daily data, adding cost and complexity.
- Hard-coding credentials or connection details in pipeline code.
- Skipping tests because pipelines are seen as plumbing rather than software.
- Allowing dozens of near-duplicate pipelines to extract the same source data.
Frequently asked questions
Is ETL obsolete?
No. ETL remains the right choice when sensitive data must be filtered before loading or when targets have limited compute.
When is streaming worth the effort?
When decisions lose value within minutes, such as fraud detection, operational alerts or real-time personalisation.
How do we handle schema changes?
Use schema evolution features where available, validate schemas at ingestion and communicate changes through data contracts.
A 90-day action plan
Days 1 to 30: inventory existing jobs, their owners, schedules and failure rates, and identify the reports that suffer most from late or wrong numbers.
Days 31 to 60: move the most troublesome flows into version control with automated tests, add freshness and volume alerts, and document lineage for key outputs.
Days 61 to 90: introduce data contracts with two source teams, retire duplicate extracts and measure the reduction in incidents and manual fixes.
Questions to ask before choosing tools
- Which sources need prebuilt connectors, and which need custom code?
- Do transformations need to run inside the warehouse, in Spark or both?
- How will we test, review and deploy changes?
- How are failures retried, and how are operators alerted?
- What are the licence and compute costs at our expected volumes?
Key terms explained
- Change data capture: reading inserts, updates and deletes from a database log in near real time.
- Orchestrator: a tool that schedules jobs and manages dependencies between them.
- Backfill: reprocessing historical periods after a fix or change.
- Data contract: an agreement on schema, meaning and quality between producers and consumers.
- Idempotence: the property that rerunning a job gives the same result.
The bottom line
Reliable movement of information is the unglamorous foundation of every analytics and AI initiative. Choosing the right pattern for each use case, treating pipeline code like software, agreeing contracts with source teams and monitoring freshness and volume all reduce incidents and rework. Avoid unnecessary complexity, retire duplicate extracts and document lineage. Well-engineered flows give analysts and business leaders numbers they can trust.
Further reading on data pipelines
For authoritative, vendor-neutral guidance on data pipelines, see the Apache Kafka project. You can also browse our free whitepapers.

