HomeExpertise03 · AI integration

AI integration

RAG pipelines, embeddings and vector search with pgvector — in production use, not as a demo. This page shows what it fails on, how I built it, and what you can check it against.

The work in progress: 38 PDFs get chunked, embedded — and one question gets an answer with a sourceAI-generated
The work, running — the only question anyone arrives here with.21:9
Area
03 of 04
Stack
11 technologies
Reference
AI profile generator
What the work leaves behind: the chunks table with 2,412 rows, HNSW and GIN indexesAI-generated
What the work leaves behind: the table that answers.21:9

The situation: search finds no answers

The knowledge is there — in PDFs, tickets and shared drives — but nobody can find it. Search finds words, not answers; so you ask a colleague, and they ask the next one.

Inbox full of where-do-I-find-it questions: price list, contract, data sheetAI-generated
The same question, asked again every week. The answer exists — it just isn't in anyone's search index.16:9
Screencast · zero results for a document sitting one row below in the folder
Zero results, with the file visible one row below. Search compares characters, not meaning.21:9

Way of working

The corpus first, the model second: documents are cut at their headings, embedded and combined with full-text search. Testing runs against real questions from everyday work — nothing goes live below 90 % hit@3.

The method in operation: the pipeline on the left, the retrieval test with its scores on the rightAI-generated
Cut, embed, check — one pass16:9

EvidenceBefore and after2 exhibits

The thinking before the code: a hand-drawn chunking strategy on paperAI-generated
Before: cut at headings, never mid-sentencePhoto · 16:9
After: 50 everyday test questions, hit@3 at 94 percentAI-generated
After: the step that proved it16:9

The reference project for this area

A capability with no project behind it is a list of technologies. This is the project — with the number it produced, and the case study where you can check it.

Screencast: the proposal generator running — PDF preview left, log right

Recruiting · AI

3 min

per proposal · was 45

Exposés that write themselves in three minutes

Applications and CVs become a finished candidate profile — Claude writes the copy, the pipeline sets the document.

Next.js TypeScript PostgreSQL pgvector RAG Claude / GPT / Gemini AWS (S3, Lambda) Docker

Inside the pipeline: annotated code and the log of one finished proposalAI-generated
The same case from inside: the pipeline16:9
The same case from outside: the printed proposal on the steel tableAI-generated
The same case from outside: the resultPhoto · 16:9

All reference projects

What I worked on

Four kinds of work, each with the one picture that makes it checkable. Numbered so they can be pointed at — not because one follows another: each stands on its own.

  • 01

    Document processing and automated parsing

    Parser: the invoice on the left, the extracted fields with their checkmarks on the rightAI-generated
  • 02

    Semantic search across your own data

    Screencast: semantic search with its sources named
  • 03

    Chat and voice bots connected to existing systems

    An assistant wired into the systems: a stock figure with ERP named as its source, liveAI-generated
  • 04

    Replacing AI SaaS with an integrated in-house build

    The arithmetic: €1,480 of SaaS against €90 of self-hosting per monthAI-generated

The stack: pgvector, RAG, HNSW

After the situation and the way of working, nobody is still asking what is installed — they are asking whether any of it is real. The line stays; two frames below it answer.

Claude GPT Gemini pgvector PostgreSQL Python Node Docker Twilio RAG HNSW

In use

API log in operation: models, tokens, latency — with the fallback lineAI-generated

Repeatable

What makes the operation repeatable: the compose file and the CI workflowAI-generated

Scale

Small meant: a script that sorts invoice emails from Friday on. Large meant: search and assistance across 14,000 documents and three departments — the work sat in the edge cases.

Small case
Search · Over the corpusAI-generated
Exists by Friday, holds stillPhoto · 16:9
Large case
Pipeline · The special casesAI-generated
A process — a still would be one arbitrary moment16:9

Common questions

Five questions that keep coming up about this area — answered from the projects they came up in.

What is RAG, and when is it the better build than a chatbot?
RAG means: search the documents first, then let the model answer with what it found. It is the right shape as soon as the answer sits in the material itself and not in the model’s world knowledge.
Where did the documents live — in-house, or with the model provider?
The index lived in the client’s own PostgreSQL. Only the excerpt the question needs goes to the model — and where even that was too much, the model ran on their own machine.
How many documents can a pgvector search handle?
14,000 documents across three departments run in production, with HNSW as the index. In practice the limit is not the number of documents, it is how cleanly they are cut.
How was it measured whether the search is good enough?
Against real questions from everyday work, not against a demo. The test set was built from questions that were actually asked; nothing went live below 90 % hit@3.
Is a dedicated vector database like Pinecone needed?
Rarely. pgvector sits inside the database that is already running — one system less, one backup less, one bill less. Past a few hundred million vectors that answer changes.

Contact

Half a minute, and you know more.

Half a minute on what I work on and how. If a question is left, write to me — an answer within one working day.

Four areas

AI integration rarely stands alone — in most projects this area reaches into at least one of the others. That is why the chain belongs together.

One change running through four layers: interface, API, database, cloud — with a timestamp at each stationAI-generated
One change, four layers — why the areas belong together21:9
An order list and its detail view side by side: ORD-4418 selected, net, VAT and total beside itAI-generated

01

Full-stack development

React and Next.js frontends with SSR and Server Components. Backends with TypeScript, Python and PostgreSQL.

A deployment log after the push: image, push, migrations, probe, routing — live after 1:44AI-generated

02

Cloud & Deployment

AWS architecture, deployment via Docker, clean secrets management. Or your own server, when that's the better math.

Sync run 4471: 3,184 records, one mismatch — ERP 214 against shop 209, decision HOLDAI-generated

04

System integration

Middleware between systems that don't talk to each other. Marketplace connections, stock and order data synchronization.

All four areas