Weekdays and Holidays of a Data Engineer Prototyping Enterprise AI Assistants
There's a profession that's barely mentioned at conferences, described in three lines in job postings, and yet nothing launches on a project without it. It's the Data Engineer who builds prototypes of enterprise AI assistants..

Article contents8×
There's a profession that's barely mentioned at conferences, described in three lines in job postings, and yet nothing launches on a project without it. It's the Data Engineer who builds prototypes of enterprise AI assistants. They live at the intersection: preparing data for LLMs, assembling context for RAG, testing retrieval, catching hallucinations, and turning raw corporate chaos into a working assistant. Let's break down what their weekdays consist of and what counts as a holiday.
Who they are and why the role appeared right now
A classic Data Engineer builds pipelines: pulls data from sources, cleans it, puts it in storage, and provides access for analysts and ML teams. The work is clear, the tools are established, the metrics are defined. AI assistants broke that picture because they demanded a different relationship with data.
For an LLM, what matters isn't so much the storage structure as the quality of context. For RAG, what's critical is how documents are chunked, how retrieval is built, how entities are resolved. For quality evaluation, you need to measure accuracy and understand exactly where the system fails: didn't find the right document, found it but didn't use it, used it but distorted it. All of this is new work that fits poorly into classic DE responsibilities.
That's how the hybrid role emerged. It requires engineering skills and an understanding of how an LLM works, how it fails, how to test it. The combination is rare, so specialists are few, and demand is growing faster than the market can train people.
What the work actually involves
Let's start with data. An assistant needs sources — internal system APIs, databases, files, wikis, support tickets, correspondence. Each source is its own story with its own formats, access rights, and update frequency. The Data Engineer assembles all of this into a single layer usable by the model.
Classic ETL only works here up to a point. Data needs to be delivered and prepared for semantic search: documents split into fragments, connections between them preserved, metadata enriched, freshness ensured. A mistake at this stage doesn't show up immediately — it surfaces weeks later, when the assistant starts confidently answering with outdated information.
Next comes context for RAG. That's a discipline of its own. You need to figure out how to chunk documents so fragments are meaningful and self-contained. How to build an index so retrieval finds what's relevant. How to resolve entities — so that "Ivan Petrov," "I. Petrov," and "ivan.petrov@company.ru" are the same person. Mistakes here are expensive: the assistant either doesn't find what's needed or finds the wrong thing.
Then comes quality evaluation. Hallucinations are concrete cases that need to be caught and classified. Where exactly does the model fail: in retrieval, in reasoning, in answer formatting? To understand that, you need data and tools. A golden dataset — a set of reference questions and correct answers. Automated runs — mass tests showing whether things got better or worse. Retrieval and hallucination metrics — the numbers that reveal the trend.
And finally — the prototype. Assembled, working, tested. Ready to be handed off to the production team, who will take it and bring it to the client.
Weekdays: what it looks like in practice
The morning starts with logs. Overnight the assistant answered two hundred queries, some of them wrong. You need to break them down: where the failure is, why, what class it belongs to. Fifteen minutes — and the first hypothesis is ready: the problem is in chunking, documents are split so the key paragraph lands in different fragments and gets lost.
Then come the fixes. Rebuilding the index, running again, comparing metrics. Sometimes it helps, sometimes it doesn't. Then you test the retrieval hypothesis: maybe the embeddings don't distinguish the right entities because the corpus contains similar documents. Or the problem is in the prompt: the model got the right context but didn't use it.
By noon a client arrives with a new requirement: the assistant must answer not only from documents but also from CRM data. So you add a source, figure out how to connect it with the rest, update the golden dataset. Several days of work.
After lunch — a meeting with the product team. We discuss metrics: what counts as success, what thresholds are acceptable, how we measure. Everyone argues, because metrics in AI are a slippery subject. 85% accuracy sounds good, but if the remaining 15% are critical errors on financial questions, the number stops being reassuring.
The evening goes to test runs. A mass run of a hundred questions, automated scoring, failure analysis. By the end of the day — a report: where things improved, where they got worse, which hypotheses held up. And a task list for tomorrow.
Holidays: what counts as a win
A holiday in this work isn't a launch to production or a model update. A holiday is when everything comes together. When retrieval finds exactly the document that's needed. When the assistant answers precisely, without hallucinations. When the golden dataset passes at ninety-five percent, and you know it's the result of weeks of work, not chance.
Another holiday is when you manage to localize a complex bug. Not "the assistant sometimes lies," but "the assistant fails when the question contains a negation and two entity names." That kind of localization is half the solution.
And the third — when the prototype goes to production and the production team doesn't send it back a week later asking "how does this work." That means the handoff succeeded, the documentation is complete, the tests are reproducible. A rare and pleasant feeling.
Why it's harder than it looks
A classic DE project can be broken into stages and parallelized. An AI prototype can't be split that way, because everything is connected. Change the chunking — retrieval results change. Change retrieval — answers change. Change the prompt — evaluation changes. It's a tightly coupled system, and any change requires re-verification.
Uncertainty adds to it. In classic DE, "the right result" is obvious: the number matches, the report builds. In AI, "the right answer" is a fuzzy concept. Two specialists can assess the same answer differently. That's why the golden dataset and clear criteria matter so much: without them, evaluation turns into polite opinions.
And third — constant learning. Tools change every few months. What worked six months ago is outdated today. You have to read, try, re-verify. It's a mode of work, not a one-time investment.
Author's column
Climax: why the role will grow
In the end, a Data Engineer for AI prototypes is a role that emerged because AI projects don't work without it. The model is only part of the system. Everything else — data, context, evaluation, tests, metrics — is where quality is made.
As companies move from experiments to real deployments, demand for such specialists will only grow. Launching an assistant is easy. Making it answer correctly, reliably, and predictably is work that demands engineering depth, systems thinking, and patience.
Everything is within our power. But only if we honestly admit: in AI assistants, the win goes to whoever has better-prepared data and better-built context.
Glossary of terms
- Data Engineer (DE) — a specialist who collects, processes, and stores data for analytics and ML systems.
- AI assistant — an LLM-based system that answers user questions and performs tasks within a defined domain.
- Prototype — an early version of a system for validating an idea. Not intended for real users without further work.
- Production (prod) — the working environment where the system is used by real users.
- RAG (Retrieval-Augmented Generation) — an approach where the model answers based on retrieved external documents.
- Context for an LLM — the set of data and instructions fed to the model along with the question.
- LLM (Large Language Model) — a large language model, the foundation of generative AI systems.
- Hallucination — a confident but incorrect model answer not grounded in data.
- Chunking — splitting documents into fragments for indexing and retrieval.
- Embedding — a numerical representation of text that allows comparing semantic similarity.
- Semantic search — search by meaning rather than exact word matching.
- Index — a structure that speeds up document retrieval.
- Entity resolution — bringing different variants of the same object to a single form.
- Golden dataset — a set of reference questions and correct answers for evaluating assistant quality.
- Automated run — a mass launch of scenarios with automatic result scoring.
- Retrieval — the stage of finding relevant documents for an answer.
- Hallucinations — a metric reflecting the share of answers containing fabricated data.
- Quality metric — a measurable indicator of system performance: accuracy, completeness, latency, cost.
- Pipeline — an automated chain of data processing.
- API — a programmatic interface for interaction between systems.
- ETL — extract, transform, load: pulling data from sources into storage.
- Prompt — an instruction for the model defining how it should answer.
- Error localization — finding the specific cause of a failure rather than a vague "the system works badly."
- Handoff — transferring a prototype and documentation to the production team.
Respectfully,
Yuri Eliseev
AI Systems Architect · Full-Stack Product Engineer
Need a project of any complexity?
Let’s discuss an idea, product, AI system or technical challenge and define a realistic first step.
