AWS Big Data Engineer
Summary
Our teams integrate into the client's operations. As a Big Data Engineer, you will collaborate with business stakeholders and your data team to deliver a data product that is sustainable and highly maintainable long-term.
You will build solutions using big data technologies such as Apache Spark, Hive/Hadoop, and distributed query engines. You will work in a large, extremely complex and dynamic data environment. You will design, develop, and operate high-scalable, high-performance, low-cost data pipelines in distributed data processing platforms. You will integrate diverse data sources, including real-time, streaming, batch, and API-based data, to enrich platform insights and drive data-driven decision-making.
What you'll do (roles & responsibilities)
- Work with product and program managers, engineering leaders, and business leaders to build data architectures and platforms to support business objectives
- Design, develop, and operate high-scalable, high-performance, low-cost, and accurate data pipelines in distributed data processing platforms
- Design, build and own all components of a high-volume data warehouse end to end
- Provide end-to-end data engineering support for project lifecycle execution (design, execution and risk assessment)
- Implement big data solutions for distributed computing
- Ensure proper data governance policies are followed by implementing or validating data lineage, quality checks, classification, etc.
- Integrate diverse data sources, including real-time, streaming, batch, and API-based data
- Interface with other technology teams to extract, transform, and load (ETL) data from a wide variety of data sources
- Design and manage data orchestration workflows using tools such as Apache NiFi, Apache Airflow, or similar platforms
- Keep up to date with big data technologies, evaluate and make decisions around the use of new or existing software products
- Recognize and adopt best practices in data processing, reporting, and analysis: data integrity, test design, analysis, validation, and documentation
- Continually improve ongoing reporting and analysis processes, automating or simplifying self-service support for customers
- Own the functional and nonfunctional scaling of software systems in your ownership area
What you should have (Must-Haves)
- Bachelor's degree required (master's an asset) in Software Engineering, Computer Science, or equivalent work experience in a technology or business environment
- Must-have certifications: AWS Big Data Specialty, AWS Solution Architect Associate, AWS Solution Architect Professional
- Minimum of 7 years of experience working in structured data engineering work processes
- Minimum of 4 years of big data engineering experience
- Minimum of 4 years of experience in Data Solutions Architecture
- Minimum of 4 years of experience in integration solutions development with pipeline tools (Qlik, Talend, Informatica, DataStage, SSIS)
- Minimum of 4 years of experience with AWS technologies like Redshift, S3, RDS, EC2, Lambda, DynamoDB, Athena, AWS Glue, EMR, Kinesis, FireHose, and IAM roles and permissions
- Minimum of 4 years of SQL experience
- Experience with data modeling, warehousing and building ETL pipelines
- Proficiency in application development frameworks (Python, Java/Scala) and data processing/storage frameworks (Databricks, Spark, Kafka)
- Experience in designing and managing data orchestration workflows using tools such as Apache NiFi, Apache Airflow, or similar platforms to automate and streamline data pipelines using various types of Databricks connectors like Lakeflow and Spark Python Data Source API and Auto loader
- Experience with performance tuning of database schemas, databases, SQL, ETL jobs, and related scripts, Unity catalog and Delta Live tables
- Experience with non-relational databases / data stores (object storage, document or key-value stores, graph databases, column-family databases)
- Experience working with DevOps pipelines (Git, Gitlab, Jenkins), CI/CD, automated testing (unit, functional, performance)
- Ability to take initiative, operate independently, and drive complex data management projects to completion
- Excellent analytical and problem-solving skills, with the ability to analyze complex data issues and develop practical solutions
- Strong communication and interpersonal skills, with the ability to collaborate effectively across cross-functional teams and stakeholders
- Proficiency in English both written and spoken
Nice-to-Haves
- Experience with project management and progress tracking tools
- Experience with Agile development methodologies
- Strong knowledge of data governance frameworks and practices (e.g., DAMA, RDM, MDM, Data Quality)
- Strong knowledge of master data management and metadata management
Work details
- Location: Remote (EST timezone for all team members)
- Reporting relationship: You will report to the Product Manager
Compensation
Competitive contract rates, adjusted to your local market and country of residence. Compensation is location-based and reflects regional market standards. We benchmark against leading technology companies in your country to ensure competitive, fair pay.
What happens after you apply
Timeline: We aim to complete our process within 2-3 weeks of application.