I tell it exactly what I am looking for. It searches Kleinanzeigen, reads every listing that comes back, and gives each one a score out of 100. I look at the ones above 70.

Here is one of them: 45 euros for a machine that normally sells for 80 to 100. A TI-Nspire CX II-T CAS graphing calculator. New it is around 170 euros; used, about 105.

This is a boring object. It is also one of the best possible objects to hunt, for a reason that has nothing to do with calculators: nobody who buys one chose it. Schools and universities prescribe the exact model. If your course requires that machine, a cheaper one that does identical mathematics is not an option.

So every year there are people who need one specific thing at one specific time — and people who just finished and have that exact thing sitting in a drawer.

What I show

The problem with used marketplaces is not that there are no good deals. It is that good deals last about twenty minutes. You have three options: refresh manually, which costs you your day; accept that you will miss most of it; or write something that watches for you.

I wrote the third one, and that comes with saying so plainly: this is one platform, at low frequency, for one user, namely me.

What I am showing is the architecture — how the scheduling works, how the filtering works, how the scoring works, and how it reaches my phone. What I am not showing is how to get around anything. That part stays out of the video and out of the written guide.

If you build something like this: read the terms of whatever you point it at, first. They are all different, and some are a great deal stricter than this one.

Phase 1 · Server

PostgreSQL, Node, a process manager, nginx. Claude Code does the setup, I watch — in this case in auto mode, so there is no confirmation prompt at every step.

The scoring calls api.deepseek.com rather than a Claude key. That is a cost decision, not a quality judgement, and it has consequences later — see mistake three.

One blocker came up straight away: the chosen domain still resolved to Vercel, and certbot cannot issue a certificate until the A records point at the server. So: repoint DNS first, www as a CNAME onto the apex, let it propagate, then the certificate.

The process layout

This is the part I want to point at. Three separate processes:

ProcessPortMemory limit
WebSocket server3002200 MB
Web app3003500 MB
Scraper worker—1 GB

They are not one process, and the start order matters: the WebSocket has to be up before the scraper tries to connect to it. There is a comment in my own config file saying exactly that in capital letters, because I got it wrong often enough to write it down.

The scraper gets the most memory because it holds parsed HTML while it works through a batch.

Phase 2 · The hunt

This is the part people get wrong, and it has nothing to do with AI. A target product is not a search term, it is four things:

  1. The name. TI-Nspire CX II-T CAS. Not "calculator", not "graphing calculator". The exact model, because the exact model is what is prescribed.
  2. Required keywords. Words that must appear, or it is not the thing.
  3. Excluded keywords. The important one.
  4. Radius. How far I am willing to drive.

On the exclusions: search a used marketplace for a calculator and most of what comes back is cases, bags, spare parts and units sold as defective. Those are not near misses, they are a different product category that happens to share the name. So the exclusion list held Tasche (bag), Hülle (case), defekt.

The Tasche bug

And that put me straight into the mistake that cost me the most time.

The German word for the device is Taschenrechner — literally "pocket calculator". On the exclusion list sat Tasche. A substring match against my own exclusion list therefore silently threw away every single listing for the exact product I was hunting. The filter designed to remove accessories was removing the thing itself.

The fix is matching on word boundaries instead of substrings. There is a comment above that function in my repository with exactly those two words in it, because I never want to debug that again.

The generalised form: German builds compound nouns the way other languages build sentences. Filter German text by substring and you will eventually delete your own product.

Phase 3 · Collecting

Every fifteen minutes, at most five search links in parallel.

The first thing in the job is a mutex — a boolean and a timestamp. If the previous run is still going, this one logs how long the old one has been running, and exits. It does not queue, it does not run alongside, it skips.

That sounds obvious. Written down, it was not obvious at three in the morning when two overlapping jobs were both writing the same listings and the logs made no sense at all.

Then deduplication, and it is a batch operation on purpose. Every listing has an external id. Before anything else happens, the job pulls every id it already knows in one query, updates the last-seen timestamp on the ones that exist, and keeps only what is genuinely new.

The reason is boring and important: 50 results per search link, times five links, times 96 runs a day. Check those one at a time and you have built a denial of service against your own database.

Only the new ones go on to the detail scraper: three at a time, 500 milliseconds between batches, 30-second timeout. Those three numbers are the whole ethics of this thing. It is a budget, not a race.

Phase 4 · Filter chain

This is where I expected the model to do the work. It does not.

Stage one is deterministic. No model, no API call, no cost. Excluded keywords on word boundaries, accessory patterns (because "case for X" is not X), required keywords strictly, and a language check. Most of what comes in dies here.

There is a comment in that file I wrote for myself: this function is the main filter and it has to be strict, because no AI check follows it any more.

That "any more" is the story. I started with a language model evaluating every single listing. Then I moved to a cheaper model, because the volume made the bill uncomfortable. Then I added image analysis so it could judge condition from photos. And I turned that off twice.

Stage two is where the model still earns its place. Everything that survives stage one gets a structured evaluation: condition, completeness — for a calculator that means the cable, the case, the manual — whether the price makes sense for what is described, and seller signals from the detail page. Out comes a score from 0 to 100 with a reason attached.

Stage three is the price baseline. Twice a day, at 08:00 and 18:00, a separate job recalculates the average price per product from what it has actually seen. So "cheap" means cheap compared to this market this month, not compared to a number I typed in once.

Stage four is the threshold. 50 and above gets a notification, everything else lands in the feed and waits for me. 80 and up is the one I actually get out of my chair for.

Diagram · The four stages of the filter chain and what each one costs
Diagram · The four stages of the filter chain and what each one costs21:9

Phase 5 · The alert

Two channels: Telegram, because it is the one I actually read, and Web Push for the browser.

The notification carries the score, the price, the distance and the reason the model gave — not just the link. If I have to open the app to decide whether to open the app, the whole thing has failed.

The dashboard updates over WebSockets: a new listing comes in, the score gets attached, the badge changes. No polling.

And every listing keeps a price history. If the same item is relisted cheaper three days later, that is visible — and that is usually the moment to message the seller.

What it is for

I built it to flip calculators. Then a friend wanted a Mac. I told him to get a MacBook Pro with an M1 rather than anything newer, we entered it — same four fields, no code changes — and he ended up with one that had the original packaging and everything in place.

That is the point where I stopped calling it a reselling tool. A different product, not a different problem. And the second use is the one I defend in public.

Five mistakes

One: I scored listings before I had the listings. For a while the evaluation ran before the detail page was fetched. So the model was scoring a title and a thumbnail, and then confidently telling me about the condition of an item whose description it had never seen. The fix is one commit: scrape the details first, then evaluate. The scores got noticeably better for free, and no model was involved in that improvement.

Two: the page moved under me. There is a function in this repository whose only job is to get the images out of a listing. It has four different ways of doing that: first structured data inside the gallery elements, then CSS background images, then the general product schema on the page, and if all three come back empty, the image tags and their lazy-loading attributes directly. I did not design that. It accumulated. Every layer is a morning where something that worked yesterday returned an empty array.

Three: vision disabled twice. I built image analysis so the scoring could look at photos and judge condition — on a used calculator that is mostly the screen and the keypad, so it seemed worth having. Then I switched provider for cost reasons, and the new one does not take image input. There are two commits weeks apart, both turning the same thing off. The second exists because I had forgotten the first.

Four: port collisions. Five commits with the same error in them. A more robust PM2 restart. Port cleanup in the deploy script. Retry logic in the WebSocket server. Graceful shutdown. Sequential service start. That is not one bug — that is one bug I fixed four times badly before fixing it properly. The actual fix was accepting that three processes need a defined start order and a kill timeout. Not clever retry logic.

Five: the build tool — and a committed `.env`. Turbopack, four production builds with standalone output, then disabled. And a .env file checked into the repository. There is not much to say about that beyond the fact that it happened.

Three rules

One: deterministic first, model second. Every filter you can write as a rule, write as a rule. The model gets what is left. It is cheaper, it is faster, it is testable — and when it goes wrong, you can read the line that did it. A regular expression has never hallucinated a reason. It will happily eat your entire product category, but it will do it in a way you can find.

Two: assume the page will change. If your input comes from someone else's website, it is not an API and it will not warn you. Write the extraction as a chain of fallbacks from the start — and make an empty result loud. Silently returned empty arrays are the worst failure mode there is, because everything keeps running and the quality just quietly drops. Four fallback levels for one image is not over-engineering; it is what a year of operation looks like.

Three: every automated fetch needs a budget. Concurrency limits, delays between batches, a scheduling interval, and a lock so runs cannot overlap. Not because somebody forces you to, but because a system that stays quiet keeps working — and one that races gets blocked. And then you have nothing.

Those three rules are why the AI in this AI project ends up doing less than it did on day one. That is not a failure of the model. That is what the job actually needed.

Takeaway

I set out to build an AI tool and ended up building a filtering system with a language model in the middle of it, doing one job it is genuinely good at — behind rules that do not need it.

And the other half: a calculator somebody needed for two years and has not touched since is a perfectly good calculator. Somebody is starting that course in September. A five-year-old Mac is still a fast computer. None of that needs to be manufactured again; it needs to be used.

That is a search problem. And search problems are exactly the kind of thing you can put on a server and forget about.