Services / AI / Knowledge search

Ask your own documents a question, and get the paragraph back

Your organisation already knows the answer. It is in a policy nobody can find, a manual on a shared drive, a decision buried in a two-year-old email thread. Search that understands the question, and shows you exactly where the answer came from, turns that pile back into something usable.

Why it matters

The cost is the wait

Every organisation past about thirty people has the same quiet problem: the knowledge exists, and finding it costs more than it should. People ask a colleague instead of searching, because asking works and searching does not. The result is a handful of experienced people acting as a human index: interrupted all day, unable to take a fortnight off without the queue backing up, and carrying in their heads things that ought to be written down.

Keyword search fails for a reason that is obvious once you see it: people do not think in keywords. Somebody types "can I claim a taxi home after a late shift" and the document they need is called "Travel & Subsistence Policy v4 (final) (updated)". The fix is not a better filing system. Organisations have been trying that for thirty years. It is search that matches meaning as well as words, and then hands back the paragraph with the answer in it rather than nine documents to open.

We build the whole thing: connectors into wherever your material actually lives, an index that combines meaning-based and exact-match search, because pure AI search reliably fumbles part numbers and clause references, permission checks so nobody is shown a document they could not have opened themselves, and answers that cite the exact section so anyone can verify them in a click.

What you actually get

Built to be trusted

A knowledge assistant is trusted or abandoned within about a fortnight of launch. These are the things that decide which way it goes.

01

Both kinds of search, at the same time

Meaning-based search finds the policy when nobody remembers its title. Exact-match search finds part FG-2210-B and clause 14.3. Used alone each fails in a way people notice immediately, so we run both and merge the results before anything is shown.

02

Every answer carries its source

Answers cite the document and link to the exact section they came from, so a doubtful reader checks in one click. An answer nobody can trace is an answer nobody acts on, and in a regulated setting, one nobody can defend.

03

It inherits your permissions

Access is checked at question time, for the person asking, against the identity system you already run. Nothing is surfaced that they could not have opened themselves. Without this, knowledge search becomes the fastest route ever built for a confidential file to reach the wrong employee.

04

One version of the truth

When a policy is superseded, the old copy stops competing with the new one. Effective dates, document owners and supersession are part of the index rather than something the reader is expected to work out from two similar filenames.

05

It searches where your knowledge already is

SharePoint, Google Drive, Confluence, Notion, a ticketing system, a decade of Slack, a folder of PDFs on a server nobody has touched since 2019. Connectors index in place and keep up with changes, so no migration is needed before the search works.

06

It shows you what nobody can answer

The questions that return nothing useful are the single most valuable output of the system: a ranked list of the documentation your organisation is missing, ordered by how often people needed it. Most clients find that list more useful than they expected.

Where it earns its keep

Same pattern, different desks

The pattern is always a small number of experts being interrupted for things that are already written down. What differs is where the writing lives and what it costs when the answer is wrong.

Manufacturing & field engineering

01 · Manufacturing & field engineering

Twelve thousand pages of manuals, on a phone in a plant room

The problem
An engineer on site needs a torque setting, a fault code, a spare part number for a machine installed eleven years ago. The manual is a PDF on a shared drive, the amendment that matters is in an email, and the one person who knows is on holiday. The alternative to finding it is a second visit.
What we build
Manuals, schematics, service bulletins and past job reports indexed together and searchable in plain language from a phone, with exact matching on part and fault codes, so a query like "E14 fault on the older line" returns the right page rather than something adjacent to it.
What changes
More jobs fixed on the first visit because the answer is available at the machine, and the knowledge of your longest-serving engineers stops being a single point of failure the week they retire.
Regulated financial services

02 · Regulated financial services

Which rule applied on the day the decision was made

The problem
Compliance questions arrive with a deadline attached: what did our policy say in March, which version of the procedure was live, where is the evidence we followed it. Reconstructing that from a document library takes days, and the answer has to stand up to somebody else's scrutiny.
What we build
Policies, procedures, regulatory correspondence and committee minutes indexed with effective dates and version history, so questions can be asked about a point in time. Every answer returns the document, the version and the dates it applied between.
What changes
Evidence-gathering that used to take days becomes an afternoon, and answers arrive with citations a reviewer or a regulator can follow without having to take anyone's word for it.
Multi-site hospitality & retail operations

03 · Multi-site hospitality & retail operations

The answer a duty manager needs at 7pm on a Saturday

The problem
Standard operating procedures, allergen information, opening checklists, refund rules, what to do when the card machine dies. It is all documented somewhere central, and the person who needs it is on a busy floor with a phone and about ninety seconds.
What we build
Search that answers in one short paragraph with the source attached, works properly on a phone, and covers the awkward operational questions, with hard rules around anything safety-related, where it returns the official wording verbatim instead of a summary of it.
What changes
Consistent answers across every site, fewer escalations to area managers for things already written down, and a clear view of which procedures are read constantly and which nobody can find.

The technology

The tools behind it, named

Search quality is an engineering problem long before it is a model problem. Almost all of the difference between a useful answer and an irritating one is made in the indexing and the ranking, not in which model writes the final sentence.

6 layers · 30 technologies

01

Search engines

The exact-match half of the job, the part that makes product codes, surnames, clause numbers and anything with a hyphen in it work properly.

  • Elasticsearch
  • OpenSearch
  • Meilisearch
  • Algolia

02

Meaning-based retrieval

Vector indexes hold the meaning of each passage, so a question phrased in nobody's official vocabulary still finds the paragraph that answers it.

  • Qdrant
  • PostgreSQL + pgvector
  • LlamaIndex
  • LangChain
  • Redis

03

The models that write the answer

Retrieval finds the evidence; the model turns it into a sentence, and refuses when the evidence retrieved does not actually cover the question.

  • Anthropic Claude
  • OpenAI
  • Google Gemini
  • Mistral
  • Hugging Face
  • Ollama

04

Where your knowledge lives

Connectors that index in place and keep up as things change, so nobody has to migrate a decade of documents before the search is useful.

  • Google Drive
  • Confluence
  • Notion
  • Slack
  • Jira
  • SharePoint

05

Who is allowed to see what

Permissions checked at question time against the identity system you already run, so search inherits your existing rules rather than inventing a second set nobody maintains.

  • Microsoft Entra ID
  • Okta
  • Docker
  • Microsoft Azure
  • Cloudflare

06

Knowing whether it works

Search is judged on the questions it fails, so we track them: what was asked, what came back, whether it helped, and which topics keep returning nothing.

  • Langfuse
  • PostHog
  • Grafana
  • Sentry

Product names and logos are the property of their respective owners and are shown to describe the technologies we work with. Their use does not imply any partnership, sponsorship or endorsement.

How we deliver it

Live behind a human first

Six to ten weeks, and the first few are mostly about your content rather than our software. That is not a delay. It is where the result is decided.

01

We find out what people actually ask

A week of collecting real questions from the help desk, from Slack, from the people who get interrupted most. Search built around imagined questions serves imagined needs, and we have seen enough of those.

02

We look honestly at your content

Duplicates, three versions of the same policy, documents nobody has owned since 2019, two procedures that contradict each other. We report what we find and you decide what gets retired. Least glamorous part of the project; largest effect on the result.

03

We index in place

Connectors pull from wherever the material already lives, split documents into passages that keep their context and their headings, and record who is allowed to see each one. Nothing has to be moved into a new system first.

04

We tune retrieval before any model speaks

Your real question set is run against the index and scored on one thing: was the correct passage in the top few results. Chunking, ranking and the balance of exact and meaning-based search are adjusted until it is. Before a model is asked to write a single word.

05

Then we let it answer

Only once retrieval is good does a model turn the evidence into a sentence, with citations, and with explicit instructions to say it does not know when the retrieved material does not cover the question.

06

We watch the failures, monthly

The questions that returned nothing, the answers people rejected, the documents that turned out to be wrong. The output is a short list of things to fix: some in the search, and quite a few in your documentation.

Before you commit

The questions worth asking

Is this not just the chatbot again?

The technology overlaps; the design does not. A customer assistant handles a known set of external questions safely and in your brand's voice. Knowledge search serves staff asking anything at all, where permissions, versions and citations matter far more than tone, and where a wrong answer costs a bad decision rather than a bad review. Plenty of organisations end up with both, sharing one index.

Our documents are a mess. Do we have to fix that first?

Partly, and we would rather be blunt about it early. Search copes well with disorganised filing; that is exactly what it is for. What it cannot do is resolve a contradiction between two documents that both claim to be current, or invent a policy that was never written down. Our indexing pass surfaces both as a list, and clearing the worst of it is usually a few days of somebody's time with a genuine payoff.

Will it show people things they should not see?

Not if it is built properly. Permission is evaluated at question time for the person asking, not baked in when the document was indexed, so revoking someone's access takes effect immediately rather than at the next re-index. We test this explicitly before launch, using accounts that should not be able to reach particular documents, and we show you the results.

How is this different from the search already in SharePoint or Drive?

Those search filenames and words well, meaning poorly, and they stop at a list of links. The difference people notice is being handed the paragraph containing the answer, drawn from all of your systems at once. If your material genuinely lives in one place and your people already search in precise keywords, the built-in search may be enough, and we will say so rather than sell you something.

What about exact things: part numbers, clause references, names?

This is where pure AI search embarrasses itself, so we never rely on it alone. A traditional exact-match index runs alongside the meaning-based one and results are merged, which is why a question mixing a part code with a plain-English description returns the right page instead of something that merely reads similarly.

How do we know it is telling the truth?

Every answer carries the passage it came from, so checking takes one click, and that alone changes behaviour, because people do verify the answers that matter. Beyond that we score it against a set of questions with known correct answers and track the answers people reject. It will still be wrong occasionally, and in our experience the most common reason is that two of your own documents disagree. The citation is what makes that visible rather than invisible.

Stop being your company's search engine

Tell us the five questions your most experienced people get asked over and over. We will tell you whether search fixes it, whether documentation fixes it, and what each would take.

Start the conversation