Back

AI Engineer - Core

Hilbert's AI·LinkedIn·Turkey·Remote·Full-time·12 hours ago
AI/ML

About the role

Hilbert is building a reasoning engine that must navigate non-deterministic user behavior across data silos - turning months-long decision cycles into minutes. Fully agentic by design, our demand intelligence platform doesn't just call APIs; it solves the hard problem of orchestrating multi-step inference over messy, high-stakes enterprise data where deterministic answers don't exist.

From Fortune 500 enterprises to beloved brands like FreshDirect, Blank Street, and Levain Bakery, operators run their growth on Hilbert. We're also co-building alongside leading AI companies.

We're looking for an AI Engineer who can build production-grade AI systems end-to-end - from prototype to pipeline to product - with the ownership and urgency of a startup culture.

This is not a "wire up a prompt chain and move on" role. You'll own core pieces of the AI stack that power Hilbert's demand intelligence platform - designing agent architectures, building evaluation systems, and making hard tradeoffs between accuracy, latency, and cost in production. You'll ship fast in conditions where the spec is evolving, and communicate what you're building (and why) with clarity to the rest of the team. If you think in systems, have opinions about how agentic workflows should actually work, and want to build AI products that drive real enterprise outcomes, we want to meet you.

What you'll own first

Evaluation and testing for our agents. Hilbert's agents are in production with enterprise customers today. Before we expand what they do, we need to know, reproducibly, when a change makes them better or worse. You'll own designing the eval harness, defining what "correct" means for a multi-step agent trajectory, building the regression gates that run before anything ships, and turning production failures into test cases. From there, the scope widens into retrieval, orchestration, and execution across the AI stack.

What you'll do

  • Own the evaluation layer for our agents - harnesses, metrics, golden datasets, reg…
View all