Founding AI Engineer @ Embedded AI Efficiency Layer, Berlin
About the role
About the Venture
About the Role
What You’ll Do
- Research, invent and productionise systems for AI cost efficiency, context optimisation and token reduction.
- Build across context selection, compression, caching, model routing, tool-call reduction, retry control and output budgets.
- Develop APIs, an OpenAI-compatible gateway and integrations with model providers and enterprise AI applications.
- Create reproducible cost-quality evaluations, regression tests and quality gates.
- Own production reliability across monitoring, deployments, incidents, capacity, availability and latency.
- Build secure enterprise infrastructure with tenant isolation, access controls, secrets management and protected customer data.
- Optimise distributed-system performance across throughput, caching, rate limits, resource usage and failure handling.
- Translate early customer requirements into reliable product capabilities.
- Take on the broad day-to-day engineering work required to move an early product forward.
About You
You’re an experienced AI engineer who combines strong software engineering with original thinking about how production AI systems can become more efficient. You can reason from first principles about where cost and quality are lost, test new approaches rigorously and turn the strongest ideas into production-grade code.
- You have substantial experience in software engineering, backend systems and API development.
- You are highly proficient in python or TypeScript or at least some coding languages and comfortable working across them.
- You understand the foundations of modern LLM systems, including tokenisation, context management, inference behaviour, model evaluation and the trade-offs between cost, latency and quality.
- You have hands-on experience building with LLM APIs and understand the behaviour, cost and reliability challenges of production AI systems.
- You have practical experience designing experiments and evaluations that measure both model quality and system performance.
- You have worked with cloud infrastructure, PostgreSQL and Docker. Kubernetes experience is a strong advantage.
- Experience with vLLM, SGLang, LiteLLM or similar AI infrastructure is highly relevant.
- You understand distributed-system performance, including latency, throughput, caching, rate limits and failure handling.
- You can independently research difficult problems, test hypotheses and translate technical ideas into working systems.
- You write clear, reliable code and are comfortable owning systems through deployment and production.
- Fluent English is required. German is an advantage.
We do not expect you to have worked on every system or technology listed above. We do expect deep experience in some of these areas, strong software engineering fundamentals and the ability to develop genuine technical depth in the rest.
Why Join
Not included in the source posting: qualifications.
Skills
Who can apply
The employer didn't state any visa, work authorization, citizenship or clearance requirements in this posting. Confirm with the employer before applying.
Read automatically from the employer's posting text. Always confirm with the employer — requirements can change after a job is published.