O que a vaga pede
About the job
We are seeking a Data Engineer to design, build, and optimize modern big data pipelines and streaming solutions. This role will focus on real-time data processing, cloud-based data platforms, and scalable data architectures while partnering with data scientists and business stakeholders to support enterprise data initiatives.
Responsibilities
- Design and implement high-volume, real-time data streaming solutions using Apache Kafka and Spark Streaming
- Develop and optimize data pipelines using Spark Structured Streaming, PySpark, and Scala
- Troubleshoot, tune, and enhance Spark applications for performance and scalability
- Build, test, and maintain big data ingestion pipelines and datasets
- Deploy and support data platforms in AWS and Azure environments
- Leverage serverless technologies such as S3, Kinesis/MSK, Lambda, and Glue
- Manage Databricks environments, notebooks, Delta Lake, Delta Live Tables, and Unity Catalog
- Work with messaging platforms including Kafka, Amazon MSK, TIBCO EMS, and IBM MQ
- Ingest and process structured and semi-structured data from JSON, XML, and CSV sources
- Work with NoSQL databases and modern data storage technologies
- Develop and maintain shell scripts and data processing workflows on Unix/Linux platforms
- Collaborate with data scientists, engineers, and stakeholders to meet business data requirements
Required Skills
- Hands-on experience with Apache Kafka and Spark Streaming
- Strong experience with Spark Structured Streaming and real-time data processing
- Expertise troubleshooting and optimizing Spark applications
- Strong programming experience with Python and/or Scala (PySpark/Scala-Spark)
- Hands-on experience with Databricks
- Experience building, testing, and optimizing big data ingestion pipelines and architectures
- Experience deploying and supporting data platforms on AWS and/or Azure
- Experience with cloud-native and serverless technologies such as S3, Kinesis/MSK, Lambda, and Glue
- Strong knowledge of messaging platforms including Kafka, Amazon MSK, TIBCO EMS, or IBM MQ
- Experience managing Databricks Notebooks, Delta Lake, Delta Live Tables, and Unity Catalog
- Experience processing data from JSON, XML, and CSV formats
- Experience with NoSQL databases such as HBase and Cassandra
- Strong Unix/Linux and shell scripting experience
- Experience with data platforms including Kudu, Impala, or Delta Lake
Preferred Skills
- Experience building enterprise-scale streaming and event-driven architectures
- Experience supporting cloud-based analytics and data lake solutions
- Knowledge of modern data governance and data cataloging practices
- Experience working closely with Data Science and Analytics teams
- Exposure to large-scale distributed data processing environments