AI-assisted quantitative research · Prediction markets · Gabah Ltd

A trading system built to prove itself wrong before it trades.

Gabah is a research lab studying prediction markets — venues where a contract settles at $1 or $0 depending on whether a stated outcome happens. We build and test strategies for pricing them. Artificial intelligence works beside the researcher, reading millions of recorded events and attacking every hypothesis before capital is risked.

Venue Polymarket
Contract 5- and 15-minute
Assets BTC · ETH · SOL
Question above or below a strike at expiry

The narrow scope is deliberate. A five-minute contract resolves within five minutes, so a strategy accumulates thousands of independent, settled outcomes in weeks rather than years — the fastest honest way to find out whether an idea is real. Longer horizons, further assets and additional venues follow the same pipeline once the method has earned them.

Every strategy we built, priced at signal · $10 per trade · Apr 30 – Aug 1, 2026
kept killed
0
validation events recorded
0
strategies built across 23 families
0
hold approval right now
0
days of unbroken recording
How it works

One real market, from open to settlement

If you have not met prediction markets before, one contract explains them faster than any definition. Below is a market we recorded on 27 July; nothing is simplified or hand-picked.

1

The question

Will bitcoin be above $63,699.93 at 00:00 UTC — five minutes after the market opens.

2

Two contracts

“Yes” and “No”. At settlement one is worth exactly $1 and the other exactly $0. A “Yes” price of 0.64 means the market is putting the odds at 64% right then.

3

Settlement

There is nothing to argue about: bitcoin closed at $63,694.44 — below the strike. “No” won. The result comes from the exchange register, not from anyone’s judgement.

The “Yes” price across those five minutes
This is where we work. One minute before the close the market priced “Yes” at 0.64 — calling the outcome 64% likely. “No” won. We do not guess where bitcoin goes: we look for moments when the price drifts away from the fair probability, and take that difference. A single trade often loses; across thousands of trades the gap pays.

How big this market is

Over 94 days we recorded 103829 such markets — around 1,150 a day across three assets. A single five-minute bitcoin market turns over roughly $60,000 (median across our own markets on 27 July, from Polymarket's public API), while about $2,900 rests in the order book at any given moment. It is a liquid venue; our own limit is a different one: we enter in a narrow window near the strike, and our measured profit retention on real fills is 55–56%. Hence the honest note on capacity — returns grow with signal quality and frequency, not with the size of capital deployed.

Pre-live ledger · running now

The lab does not stop when you close the tab.

Approved strategies trade a research ledger continuously against real Polymarket order books. Below is that ledger, by coin, refreshed every five minutes. Which strategy fired stays internal — the outcome does not.

Today, by market
Last settled markets

    What this is, exactly. Our own research capital, priced at the moment each signal fired, $10 per trade — not client money and not a live brokerage statement. Execution cost is measured separately and shown further down. We publish this because a lab that only shows finished results is a marketing department.

    The research engine

    AI is not the trader here. It is the adversary.

    Every strategy on this page was proposed, attacked and mostly destroyed in a loop between a quantitative researcher and a frontier language model. The model's job is not to predict the market. Its job is to find the reason our latest idea is an illusion — before the idea meets money.

    Reads what a human cannot

    471407 recorded events across 103829 markets, each with the order book around it. The model works through them directly — cross-checking a claimed edge against every regime, session and volatility band it appears in.

    Argues against the result

    A promising backtest is treated as a suspect, not a discovery. Standard interrogation: does it survive walk-forward on months it never saw? Does it hold with the filter removed? Is the sample big enough to mean anything, or are we reading nine lucky trades?

    Finds our bugs

    The most valuable output has not been a strategy. It has been the discovery that one of our own inputs was silently broken, or that a result rested on information the strategy could not have had at the time. Both happened. Both are listed below.

    Where the human stays in charge. The model does not place orders, hold keys, or decide allocation. It reads, argues and proposes; a person deploys. Every strategy that reaches capital passed a human decision and carries a written kill-criterion.
    Falsification log

    The findings we are proudest of are the ones that cost us results.

    Every lab has these. Almost none publish them. Each entry below deleted work we had already done and numbers we would have preferred to keep — which is precisely why they belong on the front page rather than in a footnote.

    2026-07-17

    A volatility input had been six days stale

    A request for a 16-hour window was being silently truncated upstream, so one asset's volatility estimate described a period from the previous week. Every statistic that touched it — three months of it — was invalidated and the clock restarted. The asset was pulled from live trading the same day.

    3 months voided
    2026-07-17

    We were reconstructing outcomes we could simply have looked up

    Settlement results were being inferred from price rather than read from the exchange's own record. The inference was wrong on roughly one market in six, inflating measured win rate from .24 to .33 and manufacturing three separate "discoveries" that did not exist. All three were withdrawn.

    3 results withdrawn
    2026-06-26

    Paper profit was two to four times the achievable profit

    Simulated fills assumed we bought at the price the signal fired at. Re-pricing every entry against the recorded book showed the gap concentrated in exactly the trades that looked best. We built an adversarial fill model, and now report the number it produces alongside the raw one.

    model rebuilt
    2026-07-17

    A third asset looked ready and was not

    One asset had passed enough checks to be discussed as a live candidate. Measured on the exchange's real settlements rather than our reconstruction, its expected value per trade was negative. It was not deployed, and it is still not deployed.

    not deployed
    2026-06-26

    A $20,000 result that used information from the future

    One strategy family showed an exceptional return. The entry condition turned out to depend on data that only existed after the decision point — a lookahead leak. The result was struck from the record rather than quietly retired.

    struck

    Investors with data-room access can see each of these in the underlying records, including the statistics as they read before correction. We keep the wrong versions on purpose.

    Validation pipeline

    Four gates stand between an idea and a dollar.

    Ideas are cheap and backtests are generous. The pipeline exists to remove both advantages: a strategy must survive forward time it has never seen, then survive being re-priced against the book it actually would have traded.

    01
    151
    Formulated
    hypotheses built and back-tested on tick history with settlements read from the exchange
    02
    468,821
    Forward-tested
    paper trades placed in real time, no hindsight, each stored with its order book
    03
    55–56%
    Fill-audited
    of signal profit survived re-pricing against the recorded book, with latency and a price cap
    04
    25,610
    Gate-scored daily
    approved trades placed since the ledger opened, across 84 strategies
    05
    84
    Cleared for capital
    approved as of today — the set is re-scored every day and shrinks without warning
    Cumulative signal P&L: what the gate approved against what it switched off
    approved · 84 strategies switched off
    +$25,171
    approved set, signal-priced, $10 per trade
    $0.98 / $-0.36
    per trade: approved vs switched off
    71%
    of 90 days closed positive
    $1,262
    deepest drawdown in the period

    Read the headline number correctly. $25,171 is priced at the signal and does not yet subtract execution cost. Applying the 55–56% retention we measured by re-pricing entries against the recorded book, the same period is worth roughly $12,586–$14,096 at a $10 stake. Both numbers are ours; we would rather you saw them together than found the second one on your own.

    Forward model

    What a year looks like when half of it goes badly.

    Ten thousand simulated years, built by resampling measured daily results and forcing six adverse months into every one of them. Switch between raw signal pricing and the audited fill model to see how much of the outcome is execution.

    Simulated equity, 12 months · band covers 80% of paths
    Drawdown distribution

    Drawdown is set by adverse months, not by fill quality — it barely moves between the two pricing models.

    Nine recorded days: signal vs audited fills

    Measured, not simulated. The gap between the pair is the execution cost we subtract everywhere else.

    Risk framework

    The four things most likely to go wrong.

    The record is months, not years

    The gated ledger covers 90 trading days across 4 months — enough to show the gate selects, not enough to claim a cycle-tested track record. Newer strategies carry proportionally less history. Our simulation bands are wide because we will not pretend otherwise.

    Edges decay — so we rotate them

    Individual strategies stop working; that is expected, not exceptional. It is the reason we run 151 of them across 23 families rather than one. The gate re-scores everything daily and disarms without sentiment, research runs continuously, and new candidates enter the pipeline every week.

    Capacity is bounded, and that is the point

    These are micro-structure edges on small, fast markets. Returns scale with edge quality and frequency — not with capital raised. Any operator promising that more money produces proportionally more profit in this niche is describing a different business than ours.

    Losing stretches are arithmetic, not failure

    64 of 90 days closed positive, which means roughly one day in four did not, and losing weeks follow from that. Our deepest dip in the period was $1,262 against +$25,171 earned. The model additionally prices a year where the edge dies and nobody intervenes; that case is a bound, not a plan.

    FAQ

    The questions a careful investor asks second.

    Is this a fund? Are you managing my money?

    No. We are not a fund, we do not operate managed accounts, and we do not take custody of anyone's capital. Participation is equity in Gabah Ltd, an Armenian company — you own a share of the business, not a trading account. Terms are individual and documented; nothing on this page is an offer.

    What does "AI-assisted" actually mean here?

    A frontier language model works as a research partner: it reads the recorded data directly, proposes hypotheses, and — more usefully — attacks the ones we already believe. Concretely, it wrote the analysis that found the stale volatility input and the one that exposed our reconstruction error, both listed above.

    It does not hold keys, place orders, or size positions. There is no autonomous trading agent in this system, and we would not describe one as an advantage if there were.

    If it works, why do you need outside capital at all?

    Because the constraint is not trading capital — the strategies hold small positions by design and cannot absorb large size. The constraint is runway: data infrastructure, market coverage, and the researcher time that turns 151 strategies into a durable portfolio. Investment funds the lab, not the trading book.

    Your win rate is under 50%. Why is that acceptable?

    Because these are binary contracts bought below fair value. A strategy entering at $0.30 needs to be right about a third of the time to break even; ours are right 41% of the time. Win rate on its own says nothing without the entry price beside it — a system winning 90% of the time at $0.95 is losing money.

    Can I verify any of this, or do I take your word for it?

    You verify it. Data-room access comes with a personal code to the same analytics dashboard we use internally: every strategy, every trade, every day, including the losers and the withdrawn results. The numbers on this page are generated from those records automatically — we cannot quietly diverge from them.

    What happens if the whole approach stops working?

    Each strategy carries a kill-criterion and the gate disarms without discussion — that mechanism has already removed strategies worth $77,450 of losses we did not take. If the entire family of edges decayed at once, the honest answer is that the lab would need a new research direction, and our simulation prices that year explicitly rather than hiding it.

    Data room

    Look at the records before you talk to us.

    Qualified investors receive a personal access code to the live research dashboard — the working instrument, not a prepared deck. Read it first; the conversation afterwards is more useful for both of us.

    Already have an access code? Open the data room →