AI Data Engineer at Blackstone, England, £Competitive Rate (Outside IR35)

Contract Description

THE ROLE

Blackstone& are a technical consultancy who specialise in digital transformation and complex technical deliveries. 15 years ago that meant Cloud adoption. 10 years ago it was DevOps. 5 years ago, SRE and platform engineering. We still cover all of that, but right now a significant part of our work is AI adoption. And we mean actual AI adoption: in production, in Enterprises. Not just a vibe coded app on a local machine.

 

The pace and impact of each wave keeps accelerating, but we think successful adoption still comes down to two things: people and constraints. What does transformation actually look like inside the limits of an organisation, and how do we bring the people along for it?

Most AI conversations start with the model. Ours start with the data. Agents are only as good as what they can see, and right now the biggest challenge in enterprise AI is not the model choice or the prompt design. It is getting the right data into the right shape so that AI systems can actually use it. That is where you come in.

 

This is a senior role. We are looking for an AI Data Engineer who already owns the retrieval and data layer end to end and approaches every data problem through an AI lens. Someone who understands how agents consume information, why chunking strategy matters, how retrieval fails in production and how to fix it. You have built this before, you can set the standard for how it is built here, and you can bring other engineers up to it.

 

Part of this role is growing the people around you. Our delivery teams include strong data engineers who are moving into the AI data layer, and we want a senior engineer who can help them get there: pairing on real work, reviewing pipelines, and turning retrieval, chunking, evals, and governance from things people have read about into things they can build. The best measure of your impact here is not only what you ship, but how much better the engineers around you get while you do it.

 

This engagement is with one of our healthcare clients. Healthcare data is some of the most complex, sensitive, and consequential data you will encounter. It is unstructured, often fragmented across systems, and governed tightly for good reason. You do not need a background in healthcare to do this job well. You do need to be the kind of engineer who takes data quality, security, and accuracy seriously as a default, not as an afterthought.

 

WHAT YOU’LL BRING

Core Skills

  • A data engineering background with a strong, current understanding of how AI systems consume data. You know the difference between building pipelines for analytics and building pipelines for agents.
  • Hands-on experience with AWS, including the data services (S3, Glue, Lambda, RDS) and AI services (Bedrock, Knowledge Bases) and how they connect.
  • The ability to take unstructured data and transform it into structured, agent-ready formats in production, not just in a proof of concept.
  • Experience across more than one toolset and more than one company. We value engineers who have seen different ways of solving the same problems and can make informed choices as a result.
  • Confidence working with AI development tools and practices: using AI to accelerate your own engineering work, not just building systems that use AI.
  • Comfort working in a Kanban-based delivery environment: managing your own flow, picking up work continuously, and communicating progress without needing a sprint ceremony to stay on track.
  • Experience working in enterprise environments with real data governance, access controls, and compliance constraints.

 

AI Retrieval & Data Layer Expertise

This is a senior AI data engineering role, so these are skills you already have and can demonstrate, not areas you will grow into on the job:

  • Retrieval literacy. You understand how an AI agent finds and uses information, and how that goes wrong. You know retrieval at the level that lets you debug it when it fails in production.
  • Unstructured data engineering. You turn documents, not just databases, into data an agent can read: parsing, layout analysis, table extraction, and chunking strategies that reflect how information will actually be retrieved.
  • Retrieval pipelines, end to end. You build the pipeline that fetches the right material the moment an agent needs it — chunk, embed, store, retrieve — and you own every step of it.
  • Retrieval quality engineering. You make retrieval return the right answer, not just an answer: hybrid search, rerankers, contextual retrieval, and metadata filtering. You know the difference between a system that works in a demo and one that holds up in production.
  • The semantic layer. You build governed metrics so an agent reports approved numbers instead of inventing them. This matters especially in healthcare, where data accuracy is not optional.
  • Evals and golden datasets. You engineer the eval sets that define what good looks like and gate quality in CI, building the infrastructure that makes the team confident in every release.
  • Securing and governing retrieval. You enforce access at the data layer, so an agent only ever retrieves what a person is allowed to see. In a healthcare environment this is foundational, not a nice-to-have.
  • Advanced retrieval. You work with knowledge graphs for multi-hop questions, agent memory, and the interfaces that connect the retrieval layer to the agents that consume it.

 

Nice to Have

  • Databricks experience. Useful context, not a hard requirement.
  • Snowflake background. If you have worked with Snowflake and are now applying those skills in an AI context, that translates well.
  • Direct experience with AWS Bedrock Knowledge Bases and how data ingestion pipelines feed into them.
  • Prior exposure to a regulated or healthcare data environment (for example HL7 / FHIR), though a serious attitude to governance matters more than the domain.

 

DAY TO DAY

Data Engineering for AI

  • Design and build pipelines that transform unstructured enterprise data into clean, structured, agent-consumable formats.
  • Develop and refine chunking and parsing strategies that reflect how data will actually be retrieved and used by AI systems in production.
  • Build and maintain ingestion pipelines that feed into AWS Bedrock Knowledge Bases and other retrieval infrastructure.
  • Instrument the data layer so that quality, freshness, coverage, and access controls are measurable and can be monitored over time.

 

Working in the Delivery Team

  • Pick up and progress work continuously within a Kanban flow, managing your own priorities and communicating blockers early.
  • Work closely with AI engineers, architects, and delivery managers to ensure the data layer supports what the AI systems actually need.
  • Use AI-assisted development tools as a natural part of how you work, not as an afterthought.
  • Pair with and coach data engineers who are moving into the AI data layer, reviewing their work and helping them build retrieval, chunking, eval, and governance skills on real delivery.
  • Document what you build: pipeline logic, data contracts, transformation decisions, and lessons learned for the wider team.

 

Raising the Bar

  • Hold your work to high engineering standards: tested, reproducible pipelines with clear ownership and documentation.
  • Set the standard for the data layer and actively bring other engineers up to it. Growing the team's AI data engineering capability is part of the job, not a side effect of it — through pairing, code review, and sharing what good looks like.
  • Bring ideas from previous roles and toolsets. If you have seen a better way to solve a data problem elsewhere, say so.
  • Contribute to how Blackstone& thinks about the data layer in AI delivery. We are building practice here and your input matters.

 

WHO WILL THRIVE HERE

This role will suit you if you:

  • Think about data from an AI-first perspective: not just how to store and query it, but how an agent will consume it.
  • Have already owned a production retrieval and data layer, and can point to what you built and what you learned when it broke.
  • Have worked across different companies and different tooling stacks, and carry those experiences as an asset rather than a constraint.
  • Are comfortable with AI-driven development and use it naturally in your day-to-day work.
  • Thrive in a continuous flow environment. You do not need a sprint to get moving.
  • Are clear and direct when communicating with technical and non-technical colleagues alike.
  • Want ownership of a critical layer in AI delivery, not just a seat at the table.
  • Get satisfaction from levelling up the engineers around you, and see growing a strong data engineer into an AI data engineer as a genuine part of the work.
  • Are comfortable operating in a sensitive, regulated data environment. Healthcare experience is not required. A serious attitude toward data governance is.

 

ENGAGEMENT DETAILS

This engagement is offered outside IR35, reflecting the genuine independence and expertise the role requires. Key terms:

  • Rate: Competitive, reflecting the seniority and expertise of the role
  • Outside IR35 — engaged as an independent contractor via your own limited company or umbrella
  • Initial contract length: 6 months, with extensions expected subject to performance and project progression
  • Hybrid working — a mix of remote and on-site with the client as the engagement requires
  • Start date: to be agreed. We are looking to move quickly for the right person

 

Blackstone& is an equal opportunities employer. We welcome applications from all backgrounds.

www.blackstoneand.com