Job Title: Data Engineer
Location: Mclean, VA (Hybrid)
Duration: 5+ Months
Interview Process: Technical coding interview
Position Overview
We are seeking an experienced
Senior Data Engineer to support Discover Loan Integration initiatives. The ideal candidate will have strong hands-on experience building, testing, and supporting data pipelines using
Python, Apache Spark, AWS, SQL, and Airflow.
This role involves working within small, collaborative engineering pods to integrate source data into a new platform, understand data requirements, perform data mapping and validation, and develop reliable batch data pipelines. The candidate will participate in discovery activities, collaborate with cross-functional teams, and provide ongoing production support.
Key Responsibilities
- Design, develop, test, and maintain scalable data pipelines for Discover Loan Integration.
- Build and support batch data processing workflows using Python, Apache Spark, AWS, and SQL.
- Develop and orchestrate data workflows using Apache Airflow.
- Analyze source data, understand business requirements, and perform source-to-target data mapping.
- Use Databricks to connect to source systems, explore data, validate mappings, and establish baseline outputs.
- Write SQL queries to extract, transform, validate, and reconcile data.
- Identify data quality issues, investigate discrepancies, and ensure the accuracy and consistency of pipeline outputs.
- Support pipeline deployment, troubleshooting, and ongoing production operations.
- Collaborate with engineers and stakeholders in small, agile engineering pods to deliver integration initiatives.
- Participate in technical problem-solving, testing, and continuous process improvement.
Required Skills
- Strong hands-on experience with Python programming.
- Experience developing data pipelines using Apache Spark.
- Practical experience with AWS cloud services for data engineering.
- Strong SQL skills, including querying, data validation, and data analysis.
- Experience with Apache Airflow for workflow orchestration and scheduling.
- Hands-on experience building, testing, and supporting batch data pipelines.
- Strong understanding of data mapping, data transformation, source-to-target reconciliation, and data validation.
- Ability to analyze complex datasets, troubleshoot issues, and work independently.
Preferred Qualifications
- Experience with Databricks and distributed data processing.
- Previous experience in data engineering, data analysis, or production support.
- Experience developing APIs to expose or serve data.
- Experience working in Agile teams or small engineering pods.
- Strong analytical, problem-solving, and communication skills.
|