Mid leveldata

Data Engineer
Interview Questions

Covering Data Engineer interview questions — pipelines, ETL, Spark, SQL, and data architecture.. Free, no signup required.

10 questions ready

Q1
Design a data pipeline that ingests data from multiple sources (APIs, databases, files) into a data warehouse. Walk us through your architecture choices, including how you'd handle schema evolution, data quality checks, and incremental vs. full loads.
Why they ask this:* This tests your ability to design scalable, production-ready pipelines and your understanding of real-world data engineering challenges like schema management and data quality.
Q2
Explain the differences between Apache Spark and Apache Flink. When would you choose one over the other, and what are the trade-offs in terms of latency, throughput, and state management?
Why they ask this:* This assesses your knowledge of distributed processing frameworks and your ability to evaluate tools based on use cases—a critical skill for mid-level engineers.
Q3
You have a 10TB table in a data warehouse that's being queried slowly. Walk us through your approach to diagnosing the performance issue and the optimization techniques you'd consider (indexing, partitioning, materialized views, etc.).
Why they ask this:* This evaluates your practical troubleshooting skills and understanding of database optimization—essential for maintaining performant data systems.
Q4
Describe how you would implement a slowly changing dimension (SCD) in your data warehouse. What are the different types, and which would you use for customer demographic data that changes infrequently?
Q5
Tell me about a time when you discovered a critical data quality issue in a production pipeline that was affecting downstream analytics. What was the situation, what steps did you take to investigate and fix it, and how did you prevent similar issues in the future?
Q6
Describe a project where you had to work with a team member (analyst, scientist, or another engineer) who had different priorities or technical opinions than you. How did you handle the disagreement, and what was the outcome?
Q7
Give me an example of when you had to learn a new tool or technology quickly to meet a deadline. What was the situation, how did you approach the learning, and what was the result?
Q8
What would you do if you discovered that a data pipeline you built six months ago is now processing data 50% slower than when it was deployed, but the data volume hasn't increased significantly? Walk me through your diagnostic approach.
Q9
How would you handle a situation where a stakeholder requests a new data feed be added to your pipeline within 48 hours, but the source system has poor documentation and the data quality is unknown?
Q10
Imagine you're tasked with migrating a critical legacy ETL process from an on-premises SQL Server to a cloud-based solution (e.g., Snowflake + dbt). The current system has complex business logic and stakeholders are concerned about downtime. How would you plan and execute this migration?
🔒

7 questions locked

Upgrade to unlock all 10 questions with answer guides, videos & PDF

Upgrade to unlock →

Want questions tailored to a specific company?

Try the full generator →