We're ASOS, the online retailer for fashion lovers all around the world.
We exist to give our customers the confidence to be whoever they want to be, and that goes for our people too. At ASOS, you're free to be your true self without judgement, and channel your creativity into a platform used by millions.
Everyone needs some help showing up as their best self. We're Disability Confident Committed - let our Talent team know if you need any reasonable adjustments throughout the recruitment process
As a Senior AI Engineer, you will be part of the AI Platform team, helping to build and scale the shared foundations that enable AI capabilities across ASOS. The primary focus of this role will be contributing to the Agentic AI Platform initiative, alongside other core AI platform capabilities as the platform evolves.
This role is focused on the platform layer, rather than individual business use cases. You will design and implement shared standards, templates and reference implementations for agentic AI on Azure, enabling application teams to safely design, deploy and operate AI agents at enterprise scale. Working closely with Product teams, Cloud Infrastructure, Security and partners, you will help ensure AI capabilities are secure, observable, reusable and governed by default. You will also contribute to the production foundations needed to operate AI capabilities reliably, including LLMOps, model access patterns, prompt and agent lifecycle practices, observability and secure enterprise integration.
What you’ll be doing
- Designing and building AI platform capabilities on Azure, with a strong focus on agentic AI patterns such as agent runtimes, orchestration and tool integration
- Contributing to the Agentic AI Platform initiative, helping define how agents are built, integrated and operated across the organisation
- Designing and maintaining standardised templates and reference implementations for LLM and Generative AI workflows, enabling teams to adopt consistent patterns for prompt design, tool calling, multi‑step agent flows, retries and failure handling
- Implementing secure, governed access patterns for LLMs and enterprise tools using APIM, platform gateways, Entra ID, RBAC and managed identities
- Contributing to LLMOps and model runtime patterns, including standard approaches for model access, routing, caching, token optimisation and cost‑aware usage controls
- Supporting lifecycle and evaluation practices for agent configurations, prompts and AI workflows, including testing, controlled change and release readiness
- Designing secure tool-access patterns for agents, including MCP/tool abstraction, credential management and enterprise API integration.
- Contributing to AgentOps and GenAIOps capabilities, including telemetry, run history, task outcomes, error analysis and feedback loops
- Contributing to reliability patterns for production AI systems, including latency monitoring, alerting, scaling considerations and operational readiness.
- Applying CI/CD and software engineering best practices to AI platform and agentic components
- Embedding observability by default, ensuring AI systems are measurable, debuggable and auditable through logs, metrics and traces
- Partnering with Cloud Infrastructure and Security teams to design secure, scalable and cost‑effective Azure environments