Senior AI Research Engineer

Worldwide | Sept. 1, 2026

Report as Closed

Company: Improving

Country: Worldwide

Type: Remote

Employment: Full-time

Description: Improving is an IT services firm focused on AI, data, and applications. We modernize legacy systems, build cloud-native platforms, and deliver future-ready solutions through collaborative, long-term partnerships that drive measurable outcomes.

At Improving South America, we provide IT services to transform the perception of IT professionals. We focus on IT consulting, software development, and agile training. 
The company promotes an exceptional work culture based on teamwork, excellence, and fun, with a focus on personal growth and shared rewards. Upon joining, the candidate will become part of a community that prioritizes open communication and strong, long-term working relationships, supported by a framework for professional development and continuous learning.
  • Intermediate-advanced or advanced English proficiency (required)
  • 4+ years of software development experience, with 2+ years deploying LLM-based systems into production
  • Solid Python skills—async, streaming, structured outputs, observability — as well as comfort with cloud infrastructure (AWS preferred)
  • Hands-on experience with agent-based patterns: tool usage, planning loops, state management, long-range execution
  • Good evaluation instincts: able to design an evaluation that truly measures what matters, and can distinguish when a result is a signal or noise
  • Comfortable reading papers and translating ideas into working code; you don’t need a Ph.D. to do this, but you do need curiosity
  • Proven track record of taking systems from prototype to production — takes ownership of reliability, latency, and cost, in addition to quality

    Signs We Look For

  • Equally averse to “launching without measuring” and “measuring forever without launching.”
  • Pragmatic researcher / curious engineer — chooses the right approach depending on the problem.
  • Strong opinions on how agents fail at the system level, not just at the model level.
  • Communicates clearly to different audiences—can explain a complicated evaluation result to a PM and a memory architecture to a researcher.
  • Raises the team’s standard in both rigor and speed; this must be a good translation—is it correct?
At Improving South America, we provide IT services to transform the perception of IT professionals. We focus on IT consulting, software development, and agile training.
The company promotes an exceptional work culture based on teamwork, excellence, and fun, with a focus on personal growth and shared rewards. Upon joining, the candidate will become part of a community that prioritizes open communication and strong, long-term working relationships, supported by a framework for professional development and continuous learning.
What You’ll Do
  • Build production-grade agentic systems while staying close to the frontier of research—the role exists at the intersection of “just get it out there” and “is this the right approach?”
  • Prototype novel agent patterns (self-evolving loops, multi-agent orchestration, memory architectures), validate them with rigorous evaluations (evals), and then consolidate them into deployed services
  • Be responsible for the end-to-end evaluation infrastructure: design golden datasets, build LLM-as-judge harnesses, run experiments, and use the results to guide technical decisions
  • Read recent papers and decide what’s worth testing — run quick ablations, discard ideas that don’t work, and bring the ones that do to production
  • Collaborate with platform engineers on deployment, latency, and cost
  • Contribute across the entire technology stack: agent design, retrieval/ranking, prompt and context engineering, tool integration (Claude Agent SDK, MCP, Bedrock)
Benefits
  • Long-term contract.
  • 100% remote.
  • Vacation and PTO
  • Possibility of receiving 2 bonuses per year.
  • 2 salary reviews per year.
  • English classes.
  • Apple equipment.
  • Online course platform
  • Budget for purchasing books.
  • Budget for purchasing work materials
  • Much more…
Preferred
  • Experience with Claude Agent SDK, MCP, or similar agent frameworks
  • Familiarity with the LLM-as-judge methodology, statistical bias correction, or evaluation platforms (LangFuse, Braintrust, AgentCore)
  • Experience in retrieval/ranking (embeddings, hybrid search, rerankers, semantic highlighting)
  • Exposure to self-improving systems, scaling of

    Apply here:

    Web: Apply here

    Emails:



Similar Python Jobs

Found 6 similar Remote jobs

Improving
AI Engineer

Improving

Remote Full-time

Worldwide

View Job →
Niuro
Technical Lead / Head of Technology

Niuro

Remote Full-time

Worldwide

$36,000 - $42,000

View Job →
Niuro
Lead Data Engineer

Niuro

Remote Full-time

Worldwide

$48,000 - $60,000

View Job →
Improving
IA Engineer RAG & Generative AI

Improving

Remote Full-time

Worldwide

View Job →
BNP Paribas Cardif
Actuarial Research Engineer

BNP Paribas Cardif

Remote Full-time

Worldwide

View Job →
PaperStreet Web Design
Applied AI Developer, APIs and Business Intelligence

PaperStreet Web Design

Remote Full-time

Worldwide

View Job →

Find High-Paying Python Developer Jobs ($80K-$200K+)

Django, Flask, FastAPI, Tornado & More | Remote & Onsite | Updated Daily | No Recruiter Spam

Join our community of over 1,000 Python developers and get instant access to the highest-paying Python jobs worldwide - from Django, Flask and FastAPI to AI and data roles. Save time with our curated job listings featuring transparent salary ranges.