← All cases and insights Knowledge base →

ARTICLE · Approach

A Consultant With a Chatbot Is Not Yet AI Consulting

Nikita Nechaev · 18.08.2026 · 7 min

Over the past year the word "AI" has appeared in what feels like every consulting proposal on the market. Faster, cheaper, "powered by artificial intelligence." Ask the clarifying question — what exactly does that mean — and you usually get a pause.

The pause is deserved. "We use AI" tells you about as much about a firm's process as "we use electricity." My partner has already taken this apart for interview analysis: behind the same words hide operations with wildly different reliability. Here I want to do the same one level up — for a consulting practice as a whole. Because "we work with AI" covers at least five different ways of running a firm, and a client should know which of the five they are paying for.

This is a personal text — written the old way, by hand, without AI this time. Not out of distrust: AI drafts our materials, and it does so remarkably well — you will see just how well below. But an opinion is the one thing you cannot delegate. We use the machine wherever it is stronger; the decision always stays with a human. This text is that kind of decision. We spent two years rebuilding our own practice around AI, stepped on most of the rakes ourselves, and I have accumulated opinions — some of them sharp.

Five levels: from a subscription to a factory

Level one — an employee with a subscription. A consultant opens a chatbot and works faster. The firm's process hasn't changed, the quality hasn't changed, and client materials travel to someone else's cloud from a personal account — often in breach of an NDA nobody paused to think about. By my estimate, most of the market sits here, big names included. At the large firms this is called a "pilot" and lives in press releases.

Level two — AI in production. Drafts, texts and slides are generated systematically. Speed rises noticeably; quality does not. The errors are the same as before — they are just phrased more smoothly now. Fluency is the great strength of generative models as writing tools and their great danger as analysis tools: the reader mistakes confidence of form for strength of evidence.

Level three — AI in research. The firm acquires data assets of its own and the tools to reach them: knowledge bases, connectors to primary sources, long autonomous research runs. For the first time, what changes is not the speed but the raw material a conclusion is made from. More on this below — I consider this level badly underrated.

Level four — AI in control. The system checks the finished material: every figure against its source, every claim against the transcripts and the financial model. This is where quality jumps — and this is the level I almost never see on the market.

Level five — the factory. The four previous levels assembled into a single process, with one non-negotiable condition: confidential data lives inside the firm's own perimeter and never leaves it. This is what we built, and below I'll tell you what it delivers — lining and all.

Hence the first rule for a client: "do you use AI" yields nothing. The question is which level the vendor is standing on.

Speed is the least interesting thing AI gives you

Let me start with an unpopular opinion. Everyone sells AI on speed, and speed is the weakest of its effects.

Yes, our system runs a research effort overnight that used to take three weeks of junior staff. Yes, our presentations are compiled by a generator rather than drawn, and the fix for "this block overlaps the takeaway line" costs one line of code and a minute of rebuild instead of half an hour of nudging rectangles. That is pleasant, it changes project economics — but by itself it does not make a conclusion more true.

The real effect of speed is a side effect, and it is about people, not the machine. When a rebuild costs minutes, the partner stops being a proofreader. I looked at how our own edits changed over the past year: there are practically no comments about layout left. All of them are about substance — "the argument is missing here," "this is not board-level," "where is this number from." There used to be no hours left for edits like these; they were eaten by the rectangles. Speed is valuable not because the material is ready sooner, but because the freed-up time flows into the part of the work that cannot be delegated. If it flows into "more projects at the same quality" instead — the firm has bought itself level two and stopped there.

The most valuable place for AI is the last read, not the first draft

The main conclusion from two years of rebuilding our practice sounds counterintuitive: we placed AI not where it creates, but where it nitpicks.

Before any board-level material goes out, we run a contradiction audit: every slide is checked against meeting transcripts, the financial model, previous versions and public sources. Findings are graded on three levels: direct contradiction, questionable, hygiene.

Two findings from real runs, anonymized. In a board deck for a large corporation the audit caught a key figure inflated roughly twofold against the transcript — an all-time cumulative number had slipped into a slot meant for the period. I remember the system's comment verbatim: "a board member with a calculator will catch this." In an industrial project the client had spent a year quoting "39 site problems"; the system opened the source spreadsheet and found nineteen rows in the register, with the total "39" typed in beside them by hand. We then split the material into a verified part and a reconstruction — each labeled as such.

Couldn't a human do this? A human could. Once, fresh, on ten pages. But a human runs out of anger by the fifth hour of proofreading, and the system doesn't. Thoroughness has stopped being a function of fatigue, and to me that is the most serious quality shift in this profession in all my years in it.

Let me be honest here: AI in control also catches our own errors, including the ones generated at level two. That is normal. What is not normal is a vendor who has only level two — and no one on catch duty.

Analysis built on raw material, not on retellings

The profession has a dirty secret nobody likes saying out loud: a large share of the "analysis" on the market is a retelling of other people's reports, which themselves retell data that is two years old. By the time a figure reaches a slide it has gone stale three times and lost its source twice. Forecasts built on the same second-hand sources predictably agree with each other more often than with reality.

The engineering answer is to go to primary sources. Corporate registries and mandatory disclosures. Public procurement — the state's demand in real time. Commercial court filings — a leading indicator of industry stress: lawsuits arrive before bankruptcies and headlines. Search queries — a barometer of live demand. Job postings — a proxy for investment activity: a company quietly hiring engineers in a region is building a plant long before the press release.

Every one of these sources is publicly available — which is why each is worth almost nothing on its own. The value appears in the combination: when the system lays procurement, hiring, litigation and search demand for one industry side by side, it sees a picture that exists in no report — because nobody has written that report yet.

The other half of the raw material is professional memory. For several years we have been building a knowledge base: around 14,000 reports from the world's leading practices, cut into 400,000-plus semantic fragments, searchable by meaning rather than by keywords. It grows daily and is verified at the door: of the hundreds of documents arriving each day, the best tenth makes it in. Any claim in our materials can be checked within minutes against what the best firms in the world have written on the subject — with the attribution right on the slide.

Data should not travel to models

Now for where a conversation about AI in consulting ought to begin — and almost never does.

When a consultant pastes your transcript into a public chatbot, your data physically leaves for someone else's infrastructure. Then come the questions they usually cannot answer: where is the file processed, what does the license say about training on it, how long is it retained. My partner put it this way: confidentiality is defined by the contract, not by the model's name. I will add: best of all, it is defined by architecture — by making the promise technically impossible to break.

Our principle: data does not travel to models — models come to the data. Everything confidential — the knowledge base, client materials, indexing, search — runs on local models and our own hardware. The reasoning layer receives depersonalized task statements from which no client can be reconstructed. Inside the perimeter, clients are isolated from each other at the access level: the key that opens the shared knowledge base physically cannot open a specific client's corpus. For a firm that sometimes has materials of competing groups on the same table, this is not hygiene — it is a condition of existence.

You don't have to take my word for it. We have published a line of free document tools on our site — meeting transcription with summaries and minutes, OCR for scans, version comparison, format conversion. All of them run on our own models and our own server: files go neither to OpenAI nor to Google, and are deleted right after processing. It is the architecture on public display. Consider it a test drive.

What AI does not cancel

So this doesn't read as an ad without a lining — three things that have not changed and, in my view, will not.

Responsibility for the conclusion stays with a human. The system proposes analytical objects — it does not establish their truth. We once had the system retract figures it had itself produced, stating that writing technical parameters without a source was a methodological error. I consider that a strong answer, not a weak one; but note — the decision about what to do next was still made by a person.

Rigor does not appear on its own. Wherever the requirement to verify is not stated explicitly, any model — and any junior — starts speaking smoothly. The difference between a checked and an unchecked deliverable is the difference between a table of primary sources and a well-written paragraph. Which is why verification in our process is not an option on request but a mandatory stage.

And finally: questions, arguments and decisions. The machine magnificently widens the option space and mercilessly checks the facts. What follows from all that for a specific owner with a specific appetite for risk is still a conversation between two people, and I see no horizon on which that changes.

What to ask a vendor instead of "do you use AI"

Three questions — in the spirit of the six questions to ask about a research budget.

First: where does my data physically go? The right answer contains the words "perimeter," "server" and "contract" — and contains no pause.

Second: what do you have in control? Ask them to show — not tell — how the process catches its own errors, on a live example. Level two pauses here; level four gets specific.

Third: where does the data come from — primary sources or other people's reports? And can every conclusion walk back to a specific fragment, page, recording?

The answers to these three questions will tell you more about "AI consulting" than any digital transformation deck.

How we work with this

Thesis Partners clients get access to the knowledge base by default, and it does not expire with the project: the same semantic search we use ourselves, the same reports, downloads included. The free tools are open to everyone, no registration. Soon we will begin publishing our own economic forecasts built on primary sources — and we will keep a public score of our hits.

And if you want to know what the factory would do for your specific problem — start with the 5-day express diagnostic: for a fixed fee we deliver a first position, 3–5 initiatives and a 12-week plan.

Shall we discuss your task?

Discuss your task →