Karakoram (Uzbekistan): Mid-Level AI Engineer
About the role
Karakoram builds products for Clients as well as for our own ventures. Karakoram is looking for a mid-level AI Engineer. This is a hands-on role. You will build and ship features that integrate large language models (LLMs) into real products used by a global audience and under constraints defined by Client budgets and timelines.
We want a software engineer first, and an AI-native engineer second. The strongest people at Karakoram learned to write and debug code the traditional way, before adding AI tools to their workflow. That order matters: it is what lets you notice when an AI-generated answer is wrong, instead of shipping it without checking.
We care about the quality of your work first and foremost. If your best AI projects are on your personal GitHub rather than in a past job, that is fine. We would rather see a strong personal project you built because you wanted to than a job title that never gave you the opportunity.
What you will do
Build and ship LLM-powered features into production, for Client and internal venture work.
Understand, identify and implement technical solutions to identified problems.
Design and run evaluations of LLM output: define what "good" means, build a test set, keep it updated, and catch regressions before users see them.
Build and maintain RAG (Retrieval-Augmented Generation) pipelines, agent workflows, and prompt systems.
Review AI-generated code and output carefully. Catch mistakes such as subtle bugs, security gaps, or architecture choices that will not scale.
Direct AI coding agents through multi-step tasks (for example: read the codebase, plan the change, write the code and run the tests rather than asking piece-meal).
Monitor features in production for quality, cost, latency (speed), and reliability.
Decide when an LLM is the right tool for a problem, and when regular code or a simpler approach is the better choice.
Work with a cross-functional and international team to deliver integrated solutions and results.
What we are looking for: core requirements
2+ years of professional software engineering experience, working with a team: working in a shared codebase using Git or GitLab, reviewing and merging other engineers' code, and collaborating with multiple people on the same project.
Meaningful, hands-on experience building with LLMs and AI systems. This can come from professional work, personal projects, or both.
Solid engineering fundamentals: you can read, debug, and reason about a real, complex codebase without depending on AI to understand it for you.
Direct experience designing and running evaluations of LLM output quality. You should be able to explain an evaluation you built: what it measured and what problems it caught.
Hands-on experience with at least one of: RAG pipelines, agent frameworks (LangChain, LlamaIndex, LangGraph, or a custom framework), or prompt engineering and fine-tuning.
AI-native in daily work: fluent with tools such as Cursor, Claude, and GitHub Copilot. Able to direct an AI agent through a multi-step task while carefully reviewing and assessing the quality of output
Sound judgment about running non-deterministic systems (systems that can produce different outputs for the same input) in production: awareness of cost, latency and risk areas
Strong communication for async work: comfortable working in Jira and Slack with a distributed team
Able to work flexibly across multiple timezones (meetings will largely be geared towards late Tashkent afternoons and early evenings. Individual, async work is flexible).
Responsive during agreed working hours, and able to track and update your own tasks without close supervision.
Professional working English: able to read technical documentation and discuss technical tradeoffs on a video call.
Additional Considerations: nice to have’s but not mandatory
A computer science or closely related field degree from a University, graduating between 2019-2021
Async-aware Python for I/O-bound, concurrent LLM workloads (Python code that efficiently handles many slow network calls to AI models at the same time).
Experience with more than one model provider (for example, Anthropic, OpenAI, etc).
Experience designing tool/function-calling systems, and familiarity with MCP (Model Context Protocol).
Multi-agent orchestration, coordinating multiple AI agents working together
Vector databases and embeddings (technology used to search and retrieve relevant information for RAG systems).
Awareness of guardrails, prompt-injection defense (protecting from malicious inputs) and production tracing/observability tools (monitoring what an AI application is doing in real time).
Show us your work
The best way to stand out is to show us what you have built. Send us links to your best AI projects (professional or personal) with a few sentences on: what you built, how you used AI, and one example of a mistake or weak output you caught and fixed. A personal GitHub with thoughtful AI-native projects tells us as much as, and often more than, a job title.
How we work
We are a distributed, async-first team. We use Jira and Slack for all projects and day-to-day communication. We work across time zones and ship real products for real clients and ventures. We value people who are responsive, self-directed, and honest about what is working and what is not.
Role details
Location: remote, based in Uzbekistan
Compensation: competitive, based on experience