About the Role
Production LLM agents for a multi-brand e-commerce group · full-time · reports to the Development Manager
— AI & Technology Transformation
The role
We are hiring the engineer who owns that ecosystem as a product. Today the agents work; your job is
to make them dependable, measured, and steadily more capable: maintain and harden what runs, build
what comes next, and give the whole fleet the engineering discipline — evaluation, monitoring,
governance — that turns working systems into compounding infrastructure. You will build on a written
platform architecture and decision records, on a single open-source agent framework, next to a cross-
brand BigQuery warehouse that serves as the company’s single source of truth.
What you will do
• Own the production agent fleet day to day — reliability, monitoring, incident triage, and
continuous improvement of agents colleagues rely on every morning. When an agent misbehaves,
you are the person who understands why.
• Build new agents and capabilities on the company’s standard open-source agent framework, from
business conversation to production: scoping with the stakeholders who will use them, prompt and
tool design, deployment, and hand-off into an owned-and-evaluated agent registry.
• Build the evaluation discipline — regression evals, output-quality scoring, per-agent cost tracking,
and failure-mode analysis, so “the agent got better” is a measured claim, not an impression.
• Integrate the agent memory layer — a self-hosted open-source memory service that lets agents
retain, recall, and reflect — honoring the platform’s one-way lineage rule: facts flow from the
BigQuery warehouse into agent memory, never the reverse.
• Migrate existing agents onto the company-run cloud runtime without breaking the live Slack
workflows their users depend on — a careful, one-agent-at-a-time program with a defined order.
• Do the integration engineering agents depend on — LLM APIs, MCP integrations, Shopify and
marketing/analytics APIs, and clean handoffs to deterministic workflow automation (n8n) — with
the judgment to know what should be agentic, what should be deterministic, and what must stay
human.
• Implement governance in code — permissions, approval gates, and human-in-the-loop routing,
including flows where content is compliance-sensitive and human approval is not optional.
• Work AI-assisted — the team already adopted Claude Code and Cursor on its own initiative; you
work that way too and help raise the shared standard.
What we are looking for
• 5+ years of software engineering with strong production Python — you have shipped and
operated systems with real users, and it shows in how you log, test, and monitor.
• Real, shipped LLM/agent experience — this is a must: you have built agents or LLM-powered
systems on model APIs (Anthropic, OpenAI, or similar) — tool use / function calling, structured
outputs, prompt design, context management — and you can reason concretely about evaluation,
cost, and failure modes.
• Integration depth: REST APIs, webhooks, third-party platform APIs; solid SQL (BigQuery a plus).
• Production ownership habits: observability, defensive error handling, graceful degradation, cost
awareness — agents fail in stranger ways than CRUD apps, and you design for that.
• Enough infrastructure fluency to deploy and operate your own work — Docker and cloud basics
(we run on GCP). A dedicated platform engineer owns the foundation; you are not helpless without
them.
• Working proficiency in English. Knowledge of Armenian is a bonus. (The engineering team is
Armenian and Russian speaking. Documentation and company leadership work in English).
• A bonus, not a requirement: agent memory systems / RAG / vector stores, open-source agent
frameworks, TypeScript/JavaScript, Shopify or DTC e-commerce context, workflow tools (n8n),
marketing and attribution data.
Why this role, why now
Most companies hiring for “AI agents” are hiring someone to find out whether the idea works. Here that
question is answered, agents run in production with named owners and daily users, the platform
architecture is written, and the partners have funded the roadmap that scales it.
What the company needs is the engineer who industrializes the proof: one framework, one source of
truth, one memory layer, every agent owned and measured. The work compounds, every agent you
harden and every eval you write makes the next one cheaper.
You will not spend your first year proving agents are worth it; you will spend it making them
dependable.
Apply to marko@apexteamhiring.co