Company: YOM
Country: Worldwide
Salary: $24,000 - $30,000
Type: Remote
Employment: Full-time
Description: We are a company that seeks to enable the growth and prosperity of neighborhood grocers through the digitalization of the commercial value chain in the traditional channel through technological solutions.
Experience 5+ years in data roles, or 3+ years as a Data Engineer building and operating pipelines in production. Python and SQL Advanced Python for data processing (pandas), with automated tests (pytest). Advanced SQL in PostgreSQL: modeling, indexes, execution plans and incremental loads (upserts). Integration Experience integrating heterogeneous sources: REST APIs (authentication, paging), webhooks, SFTP, flat files and third-party databases. Design of ETL/ELT processes with an orchestrator (ideally Airflow). Data quality and contracts Data validation with schemas and contracts (pandera, pydantic, Great Expectations or similar). Criterion to think about data quality from the pipeline design. Cloud AWS: S3, RDS and at least one processing service (Glue, Lambda, Batch or ECS). Your mission will be to design, integrate and operate our clients' data flows, ensuring quality, traceability and figures that match the origin. You will report to the Intelligence Leader. Behind every warehouse, hardware store or bottle shop there is a family that we want to help sell more and be better supplied. And it all starts with reliable and timely data. We are obsessive with precision: if the ERP records 100 sales and we see 98, we look for the 2 missing ones. Incorrect data affects business decisions and the user experience of our Conversational Agent. Today we are transforming our technology: we unify the stack, standardize integrations from ERPs, SFTP and APIs, and incorporate quality validations at each stage. It is a strategic priority of the company, with assigned resources, and we are looking for someone to lead this challenge together with the Data Intelligence team. Design, build and operate integration pipelines from the origin of each client (ERPs, REST APIs, webhooks, SFTP, VPN, files in S3) to PostgreSQL in AWS RDS. Write ETL and ELT processes in Python, orchestrated with Apache Airflow. Migrate legacy pipelines (AWS Glue and legacy flows) to the new stack. Define declarative data contracts by client and entity: schema, types, vocabularies, delivery windows and quarantine thresholds. Implement quality rules by dimension (completeness, validity, uniqueness, referential integrity, consistency, timeliness) and verify that each transformation preserves the data invariants. Design the data model, version it with migrations and onboarding new clients. Build pipeline observability: quality findings per run, alerts and detection of silent failures, such as the process that ends “successfully” without having received a single file. Square our figures with those of the client, explain each difference and fix the cause in the pipeline. Manage access and credentials of processes and service accounts with minimum privilege. Document flows, data contracts and operational procedures. Work with data science, product and engineering on the data that feeds the recommendation engine and platform. Real pipelines in production, with end-to-end ownership. A growing team, with a modern stack and intensive use of AI on a daily basis. Hybrid modality and flexible schedule. Benefits for sports, education, purchase of equipment and administrative days. 5 extra days of vacation Horizontal and close culture. You designed batch and streaming integrations (events, webhooks, CDC or message queues). You migrated pipelines between platforms without stopping the operation. You handle immigration
We are currently looking for our next Data Scientist to join the Intelligence team.
We have a hybrid modality, being able to come to our office located in Las Condes, Santiago, and work from home or another place :)
Apply here:
Web: Apply here
Emails:
Found 6 similar Remote jobs