Lead SRE - AI
Abingdon, LND, GB, OX14 4RW
We are looking for the right people — people who want to innovate, achieve, grow and lead. We attract and retain the best talent by investing in our employees and empowering them to develop themselves and their careers. Experience the challenges, rewards and opportunity of working for one of the world’s largest providers of products and services to the global energy industry.
Job Duties
About Landmark
Landmark, a Halliburton company, builds the software and data platforms that help the global energy industry make better decisions. Our products span subsurface interpretation, well construction planning, reservoir simulation, production optimization, and digital operations. These are tools used daily by engineers and scientists at the world’s largest energy companies and run as cloud-native SaaS platforms and as enterprise on-premises solutions.
About the Role
Landmark’s Site Reliability Engineering (SRE) team is looking for a hands-on AI SRE team lead to design, develop, and operationalize AI- and LLM-driven solutions that augment and automate SRE workflows at scale. You will work at the intersection of full-stack engineering, machine learning operations, and cloud platform reliability, building real tools that engineers depend on every day. As a member of the SRE team, you will lead an AI sub-team while working closely with development engineering in the DevSecOps organization, regional technical engineers, Halliburton Information Technology engineers, and architects. Bring your expertise in AI, full-stack development, and cloud operations to an environment offering broad exposure to today’s most relevant AI technologies, LLM frameworks, cloud platforms, and petrotechnical science.
Overall Responsibilities:
The AI SRE team lead is a critical and visible role, central to delivering AI-powered tooling across a multi-tiered cloud infrastructure. You will design and deliver full-stack AI solutions that reduce toil, improve incident response, and surface actionable intelligence from complex multi-cloud operational data. You will collaborate directly with SRE managers, development leaders, architects, and product owners to ship AI-first tooling that scales with the platform.
- Design and build full-stack AI-powered applications using Angular frontends, Python/FastAPI backends, and integrated LLM and agent pipelines. Develop and operationalize LLM-based solutions for SRE use cases, including log analysis, anomaly detection, runbook automation, incident triage, and RCA summarization.
- Implement RAG (Retrieval-Augmented Generation) pipelines over operational knowledge bases and runbooks.
- Integrate AI/ML models into SRE observability tools like New Relic and automation workflows, including monitoring, alerting, and self-healing.
- Build and maintain RESTful and async APIs using FastAPI for internal SRE tooling and dashboards.
- Develop Angular dashboards that present AI-generated insights to SRE, account, and sales stakeholder teams.
- Collaborate with the Cloud Automation team on AI-driven infrastructure automation using Terraform, Ansible, and Jenkins.
- Enforce engineering best practices, including code reviews, unit testing, and CI/CD pipelines for both ML and application code.
- Evaluate and introduce emerging AI frameworks and LLM providers, including OpenAI, Azure OpenAI, Bedrock, LangChain, and LlamaIndex.
- Maintain Kubernetes-deployed AI workloads and troubleshoot inference latency, scaling, and reliability issues.
- Participate in the on-call rotation and apply AI tooling to improve incident response times.
- Coach and mentor team members on AI integration patterns and best practices.
Qualifications & Experience
- Must have solid experience of AWS/Azure DevOps, Terraform and Ansible
- Bachelor’s degree in a technical field and 8–12+ years of professional experience in software or platform engineering.
- Proven hands-on full-stack development experience with Angular (TypeScript) frontends and Python backends in production environments.
- Strong Python proficiency, including FastAPI, async patterns, Pydantic, and SQLAlchemy or an equivalent ORM.
- Direct experience building and deploying LLM-integrated applications, including prompt engineering, function calling, RAG pipelines, and agent frameworks such as LangChain, LlamaIndex, AutoGen, CrewAI, or equivalent. Familiarity with LLM providers and APIs, including Azure OpenAI, AWS Bedrock, OpenAI API, or open-source models such as Llama and Mistral.
- Experience with vector databases and embedding workflows, including Chroma, Weaviate, pgvector, Azure AI Search, or equivalent.
- Solid understanding of CI/CD pipelines for application code and ML/AI artifacts, including MLflow, DVC, or equivalent.
- Strong container and orchestration skills with Docker and Kubernetes, including the ability to deploy and troubleshoot AI inference workloads.
- Public cloud experience with Azure and/or AWS, including compute, storage, networking, and managed AI services such as Azure OpenAI and SageMaker.
- Understanding of SRE/DevOps principles, including SLOs, error budgets, observability, incident management, and toil reduction.
- Experience integrating with observability stacks such as Prometheus, Grafana, ELK/OpenSearch, Datadog, or equivalent.
- Strong interpersonal and communication skills, with the ability to explain AI capabilities and limitations to non-technical stakeholders. Preferrably
Preferred:
- Experience with Microsoft Copilot Studio for building enterprise Copilot agents and connectors integrated with M365 and Azure AI services.
- Familiarity with AI agent orchestration patterns, including multi-agent systems, tool-use loops, and memory management.
- Experience with fine-tuning or PEFT (LoRA/QLoRA) of open-source LLMs for domain-specific SRE tasks.
- Knowledge of petrotechnical, energy, or industrial IoT domains.
- Experience with identity and access management (Okta, VIDM, SSO/MFA) as it relates to securing AI APIs.
- Hands-on scripting experience with Bash, PowerShell, or similar tools for SRE automation tasks.
Minimum Qualifications: Minimum qualifications may be acquired through technical schools or equivalent related experience. Candidates having qualifications that exceed the minimum job requirements will receive consideration for higher level roles given (1) their experience, (2) additional job requirements, and/or (3) business needs. Depending on education, experience, and skill level, a variety of job opportunities might be available from the Software Engineer to Senior
Halliburton is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.
Location
97 Jubilee Avenue, Milton Park, Abingdon, Oxfordshire, OX14 4RW, United Kingdom
Job Details
Requisition Number: 209992
Experience Level: Experienced Hire
Job Family: Engineering/Science/Technology
Product Service Line: Landmark Software & Services
Full Time / Part Time: Full Time
Additional Locations for this position:
Compensation Information
Compensation is competitive and commensurate with experience.
Job Segment:
Cloud, Open Source, Testing, Technology