Opportunity radar for a B2B product team

Every morning at 05:00 UTC the radar reads the whole market a B2B company sells into – 66 live sources out of 74, from its own customers' requests to tenders, regulators, competitors, research and the trade press of every industry it serves – weighs each signal against its own history and has a tray of opportunity cards ready before the working day starts, pointing at needs buyers have now and at business models still forming. At ten minutes a record, the reading alone would keep four or five analysts busy full-time.

build time
~5 weeks build · 319 commits · runs every morning · in use by the product team
first published
last updated

In one paragraph

Every market a B2B company sells into writes its needs down long before they reach a sales call, and in more places than your team can read. This radar reads the market whole. Every morning at 05:00 UTC it goes through 66 live sources – the company’s own customers beside tenders, regulators, competitors, research and the trade press of every industry it serves – and across all 74 it has read 9,114 records, with history back to 2018. On my estimate of ten minutes a record to read it, judge it and log who needs what, that is about 1,500 hours of an analyst’s work, and an ordinary week’s 1,070 new records add some 180 more: four or five analysts doing nothing else. Each signal is weighed against its own source’s history, so the morning surfaces what is new or strengthening, and related signals gather into opportunity cards that point at needs buyers have now and at business models still taking shape. By the start of the working day the cards are on the tray. I built the radar in about five weeks and 319 commits, and since mid-September the product team has worked from that tray.

0 sources
66 read every morning · the customers' voice and the market's
0 records
read and assessed · history back to 2018
~0 analyst-hours a week
an ordinary week's 1,070 records · at 10 minutes each
0 cards
opportunity cards put forward · from 4,258 signals and 812 themes

The problem in business terms

The next line of business for a B2B company is usually in writing before a customer asks for it by name – in lost-deal notes, in tenders where a buyer has committed money, in new rules, in competitors’ job postings, in research grants and in the trade press of the industries it sells into. What stands in the way is volume: over a thousand records in an ordinary week, across dozens of sources and many languages. Read by hand, the market gets read in samples, and samples miss the faint early signals that point at business models still forming. On the day this work started, the people choosing the next bet were choosing from three real cards on their hypothesis board, each written by hand.

how the market reaches the people who choose the next bet before vs after the radar
without the radar
  • The market is read in samples – a few feeds, an annual report, the loudest customer.
  • A request in the CRM, a tender abroad and a new regulation never meet in one place.
  • "Demand is growing" is a feeling, with no baseline to measure it against.
  • Reading the market spends the hours of the people who should be deciding.
  • A month later, the team cannot say what an idea stood on.
with the radar
  • 66 sources read every morning from 05:00 UTC – the customers' voice and the market's.
  • A customer's words and a market record land on one card when they state one need.
  • Each signal is a rate against its source's own history – new, strengthening, weakening – and a source with no history says so.
  • The reading is done before the reviewers arrive; their hour goes on deciding the cards.
  • Every card opens onto the records it was built from, frozen on the day it was made.

Every record, every day, against its own past

Reading a market whole starts as an engineering problem. A feed keeps only its latest few dozen items – one trade outlet produced 50 new records in a single run – so the radar reads daily, before anything scrolls away. A source brings its history when it is switched on, back to 2018 where one exists, because a signal can only be new against a past. And each new kind of source comes in through a gate measured before it opens, so the field widens and the morning queue stays worth your hour.

The second move turns volume into signal. Each record becomes a structured signal – who needs what, quoted and dated – classified as a rate against its own source’s history: NEW, STRENGTHENING, WEAKENING or CONTRADICTING. A source with no history reads UNBASELINED, never NEW, so an unarchived feed cannot pass off its whole backlog as discoveries. A card needs two independent witnesses, and a witness is whoever speaks: three press releases from one press office count once. Cards that rest only on the company’s own customers keep a lane of their own, so the reviewer sees what buyers ask for today beside what the market is starting to ask for by itself.

cards that stand only on the company's own customers measured on 12 September · 44 of 51 queue cards had no outside witness
criterion hard gate · an outside witness required two lanes · a person decides chosen
queue cards thrown out 44 of 51 · 86% none
"a customer asked and we could not" rejected first kept, in their own lane
everything stays visible to the reviewer no yes

Three sessions in the engine room

A working partner that reads the rules first

Every session on this radar is a conversation between me and Claude Code: I bring the problem, and it reads the code, queries production and ships the fix under the same rules the radar holds itself to.

Two files travel with every session. CLAUDE.md runs to 1,332 lines, and most of its rules carry the failure that produced them, so the next session inherits the reason along with the rule. The kickoff note holds the problem statement and the lines the machine stays behind: gate scores, card moves and the hypothesis sentence belong to people; an empty search is never evidence of absence; an impact figure without a quoted source is marked unverified.

The sessions treat production the way the radar treats evidence. Every query runs inside a read-only transaction that is rolled back, and every change to what the radar reads or believes lands as a migration in git, where the next session can read why.

A wider field, and a gate measured before it opened

A public tender is the strongest evidence the radar can hold: a buyer who has committed money to a need. The research behind this wave sent one agent to each of 24 candidate portals – its API, its robots file, its terms of use – with a sceptic agent re-checking every finding, and seven national portals came back worth reading. Their collectors were written in nine parallel branches, which one integrator then merged.

They stayed switched off until a gate stood in front of extraction. Every tender names a buyer, and without a gate a car rental or an order of office furniture would have walked into the market lane as an outside witness. Two agents built the gate in parallel, three more attacked it from different angles, and a sixth fixed what they found – thirteen of fourteen real problems before the first commit. Measured against notices the study had already labelled, it agreed with them as often as the study’s two labelling passes agreed with each other, and only then did I say yes.

For the product team, the field widened by six national portals in a single afternoon, and a tender now reaches a card on the strength of what it buys.

Telling new business from old, in shadow until people agree

A radar for new business has to know the old business first. Most cards in the early queue asked for things the company already sells, and a card had no way to say so. A coverage judge now checks every card against a register of the company’s capabilities – 91 rows I reviewed myself – and says whether the need is served, partly served, open, or once served and since retired. That last verdict is the one to watch: when demand grows for a capability the company once sold only as one-off projects, it is a product asking to be built. The judge runs in shadow, and the reviewer sees its verdict as a suggestion.

I label the judge on the card itself, “Right” or “Wrong…”, without seeing another reviewer’s label. Twenty labels in, agreement stood at 70%, and the direction said more than the average: five of six misses understated what the company does.

The session stopped after one fix, measured on the labelled cards, and one experiment, dropped when it changed nothing; a third guess would have been guessing. The judge stays a suggestion until my labels say otherwise.

transmission my notes on the assistant's design · 28 September

On “we already do this”, the judge knows less than my colleagues, who sell or improve those services every day. I would put the judge’s accuracy at 80% and theirs at 90–95%, so the judge should weigh less than people’s decisions.

The Head of Sales already leans on the verdict to reject faster, which is why the judge has to prove itself on my labels first.

”Sent” means on the board

Sending a card is the decision that matters most, and the radar leaves it to a person, so the button has to tell the truth. On the morning I sent the first card, the portal said “sent”, recorded my name and offered an undo – and the board stayed empty.

The cause sat outside the code, in a proxy that treated creating and reading differently, which is why every check that only read had passed. The repair put Jira first and the decision second, so the portal can say “sent” only once the card exists.

The probe that found it was harmless by design: a create request with no summary is validated in full and refused, which proves every other field of the payload while the team’s board stays clean.

For the people at the gate, “sent” became a fact they can rely on, and the decision log a record they can audit: every send, rejection and return carries the name of the person who made it.

What runs every morning

The run starts at 05:00 UTC, so the queue is ready before the working day begins. Its ten steps go in a deliberate order: the board’s decisions are read first, so the later steps see fresh labels, and the three judges run last, so a judge’s failure leaves the rest of the run intact.

from the whole market to the morning queue unattended until a person presses send
  1. ● 01
    74 sources
    66 on · 9,114 records · read from 05:00 UTC
  2. ● 02
    4,258 signals
    who needs what, quoted and dated · a rate against its source's history
  3. ● 03
    812 themes
    signals gathered by meaning · the same facts make the same themes
  4. ● 04
    the morning queue
    320 ideas · 135 put forward, 177 dropped, 8 merged · two lanes
  5. ● 05
    the board
    a named reviewer, signed in through Slack, presses send · 106 decisions so far · outcomes read back next morning
  • unattended schedule
  • single narrow key

The company’s own customers are a third of what the radar reads; the other two thirds are the market talking among itself.

who is speakingrecordswhat it tells the people choosing the next bet
the company’s own customers – CRM requests, deal and sales notes, the support desk, team chat≈3,300what buyers ask for today, and what was lost for want of it
the trade press of the industries the company sells into≈2,350where those industries are heading before it turns into a request
industry annual reports – 1,305 quoted facts from 99 reports≈1,300problems an industry has measured in itself · context, never a witness
competitors – their blogs, news and help pages≈740where rivals are placing their bets
research papers and public research grants≈590what is becoming possible, and who is paying to find out
public tenders≈420needs a buyer has already put money behind
regulators and registers – legal acts, enforcement actions≈390new obligations, and the demand they create
job postings≈30which capabilities rivals are hiring for

Every card can be read back to front, and the card page keeps that reading honest. This is the trail of a real card from 8 September, with its subject left out:

one card, traced back to its first record what the card page shows
  1. ● 01
    the card
    statement · who has the problem · rank and gate log stored on it
  2. ● 02
    its engine room
    every number frozen 8 Sep, 10:59 · prompt extract_signal v4 · model pinned
  3. ● 03
    frozen evidence · 4
    "the evidence as it stood when the card was written"
  4. ● 04
    each signal
    the quote, the organisation, the date · the prompt version that read it
  5. ● 05
    the first record
    source, link, time fetched · the stored copy beside the original
the judges, in shadow until calibrated each one reads; none gates yet
  1. ◐ 01
    coverage
    does the company already serve this, or did it once? · people label the judge
  2. ◐ 02
    corroboration
    is the need seen outside? · 10,274 checks
  3. ◐ 03
    annual reports
    did an industry report measure this problem? · match threshold tested before use

The whole archive can also be asked a question. A research assistant the team asks in Slack answers from the same store through a registry of read-only tools – search by words and by meaning, similar cards, a company’s history, competitor moves, tenders, blind spots – at most twelve calls a turn, and computes anything that needs a join inside a sandbox holding one snapshot of the data. It has answered 114 turns so far. Before it reached the team it had to pass a red-team run and a shadow run of 39 approved questions through the real model, with code checking each answer.

The reading bill, in your own numbers

What reading this market by hand would take

An ordinary week brings the radar about 1,070 new records – 2,375 arrived in the last seven days, 1,305 of them a one-off import of annual-report facts. Put in what one record costs your analyst and what their hour costs you, and see the desk it would take.

reading by hand: —a year at this pace: —left for your people: —

runs locally · 0 network calls · deterministic

The radar’s limits live in code, where a test can check them

A radar that reads customer text needs its boundaries where a test can hold them. The model inside the run gets no tools: text goes in, schema-constrained structure comes out, and a label outside its vocabulary is dropped in Python, so a hostile page can lie in its content and do nothing more. Customer text is redacted by shape – emails, phones, card numbers, IBANs, tokens – before any model sees it. The research assistant is the one place a model holds tools, and all of them only read, through a database role that sees views and nothing else; an auditor can check those permissions directly.

app/gemini.py · the single exit to the model
text
text in → schema-constrained JSON out
no tools · pinned model version
every call → llm_calls (prompt version, tokens, error)
labels outside the enum → dropped in Python

Deploys are pulled: every five minutes the instance fetches the main branch, builds, checks its health and rolls back to the last good commit if the new one fails. A rollback restores code and never a migration, so each migration has to leave the previous commit working – a line close to a hundred automated security reviews of individual changes have held.

A business-intelligence desk’s work, on a tray every morning

The reading a business-intelligence desk would spend its week on is done before the reviewers arrive, and their hour goes on the cards. In their first week with the queue, the reviewers decided cards faster than the machine produced them – 58 rejections and 27 sends, with two cards left. What changed is the conversation at the gate: a card arrives with its witnesses, its rate against a baseline and the judges’ readings beside it, so the meeting spends its time on whether to bet, and the question of whether the evidence is real arrives already answered. The reviewers also said plainly where the machine fell short. “The first fact, yes,” the Head of Sales said of a card’s evidence, “but the others often seem not quite relevant.” Two plausible causes were measured on production before anything was built, both turned out wrong, and the card now says which of its quotes attest its statement and shows the rest apart. Their verdicts go back into the next morning’s run, so each tray comes closer to what the people at the gate are looking for.

The live constraint has moved to the supply of cards, and that is where the work goes now – more witnesses from outside the company, and judges that earn their way from shadow into the gates on the reviewers’ labels. The purpose stays the same: an ordinary week brings in about 1,070 records, in September about three cards a day came out of them, and the cards worth most arrive before any customer asks – a need the market has started to put money behind, or growing demand for a capability the company once sold only as one-off projects, which is a product asking to be built.

Five weeks. One builder. 66 sources read every morning, some 180 analyst-hours of reading a week done before the working day starts, and 135 opportunity cards put forward – each with its evidence attached.


This case describes architecture and operating discipline. The client, its industry, its product names and every source that would point at either are deliberately left abstract; what carries over to any company is the shape of the radar – read the market whole and every day, weigh each signal against its own past, keep the evidence, and leave the bet to a person.