Bongard-miniMachine intuition.
An open-weight, Jev-like decision model built on T5Gemma 2. Give it a situation and candidate answers; get probabilities without generating text.
Read once. Judge in parallel.
Fig. 1 — One object, three shadows
Square · Circle · Triangle
One situation. Different judgments.
01 — Overview
See the relation before you can explain it.
Bongard-mini is an open judgment model. You give it a situation — a message, a document, a web page, a table or an agent’s current state — together with your questions and their possible answers. It reads the situation once and returns a probability for every answer. It writes no text.
Where it fits
Routing and triage
Send each ticket, email or request to the right team, model or tool.
Screening and guardrails
Flag phishing, jailbreak attempts and risky actions before they reach a person or an agent.
Evidence and relevance
Pick the passage that answers a question, and check whether the sources support a claim.
Agent decisions
Judge an agent’s next step from its current state, with one fast call per step.
Open weights · 7.5 billion parameters · text and images · tested in 51 languages · tested on documents of up to 30,000 tokens
An experienced player sees a promising move. A reader catches irony. Each judgment draws on relationships learned through experience, often before the person can explain each step.
Intuition can use patterns that are hard to express as a list of reasons. Its value extends beyond speed. Bongard treats this kind of judgment as an independent capability to design and train.
Look at the two groups. What relation separates them? Try to see it before you reveal the rule.
The model takes its name from Mikhail Bongard’s visual puzzles, first published in 1967. Each group has six figures. A rule holds for one group and fails for the other.
LeftConvex figures
RightConcave figures
02 — Demos
One model, different situations.
Recorded responses from the public demo. Each tab replays one input and the probabilities the model returned for every candidate answer.
Choose an example, edit the input and the questions under “Questions & settings”, then click Run.
Hosted on Hugging Face ZeroGPU. Anonymous quota is limited; signing in provides more quota. You can also run the weights locally.
Recorded examples · 30 September 2026
Input
channelemail
planpro
subjectCharged twice for order #4821
bodyHi, my card was billed twice for the same order this morning. Please refund the duplicate charge.
Replaying one reading of the inputRecorded output · 5 answers, one forward pass
- Choice
Which team should handle this?
billing 0.96account 0.03technical 0.00
- Noul
Is the customer asking for a refund?
yes 0.93no 0.07
- Score
How urgent is it, from 1 to 5?
2 0.35expected 2.1
- Choice
What is the tone of the message?
calm 0.75frustrated 0.16angry 0.09
- Noul
Should a person approve the reply?
yes 0.56no 0.44
Request and response
{
"request": {
"state": {
"channel": "email",
"plan": "pro",
"subject": "Charged twice for order #4821",
"body": "Hi, my card was billed twice for the same order this morning. Please refund the duplicate charge."
},
"questions": {
"route": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": null,
"technical": null,
"shipping": null,
"account": null
}
},
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is it, from 1 to 5?",
"criteria": [
"1",
"2",
"3",
"4",
"5"
]
},
"tone": {
"type": "choice",
"instructions": "What is the tone of the message?",
"criteria": {
"calm": null,
"frustrated": null,
"angry": null
}
},
"review": {
"type": "noul",
"instructions": "Should a person approve the reply?"
}
}
},
"response": {
"answers": {
"route": {
"type": "choice",
"probs": {
"billing": 0.965,
"technical": 0.004,
"shipping": 0.001,
"account": 0.029
}
},
"refund": {
"type": "noul",
"p": 0.926
},
"urgency": {
"type": "score",
"probs": [
0.327,
0.349,
0.227,
0.077,
0.021
]
},
"tone": {
"type": "choice",
"probs": {
"calm": 0.748,
"frustrated": 0.163,
"angry": 0.089
}
},
"review": {
"type": "noul",
"p": 0.565
}
}
}
}Input
FromNorthbank Security <alerts@northbank-verify.co>
SubjectUrgent: your account has been locked
We detected unusual sign-in activity on your account. For your protection it has been locked. Verify your identity within 24 hours at http://northbank-verify.co/login, or your account will be permanently closed.
Northbank Security Team
Replaying one reading of the inputRecorded output · 5 answers, one forward pass
- Noul
Is this email a phishing attempt?
yes 0.74no 0.26
- Choice
What does the email ask the reader to do?
sign in on link 0.73reply with details 0.20nothing 0.04
- Score
How much time pressure does the email put on the reader?
Strong 0.27expected 3.3
- Noul
Is the sender's domain the bank's own domain?
no 0.81yes 0.19
- Choice
What should the mail filter do with it?
quarantine 0.48warn 0.34deliver 0.18
Request and response
{
"request": {
"state": "From: Northbank Security <alerts@northbank-verify.co>\nSubject: Urgent: your account has been locked\n\nWe detected unusual sign-in activity on your account. For your protection it has been locked. Verify your identity within 24 hours at http://northbank-verify.co/login, or your account will be permanently closed.\n\nNorthbank Security Team",
"questions": {
"phishing": {
"type": "noul",
"instructions": "Is this email a phishing attempt?"
},
"request": {
"type": "choice",
"instructions": "What does the email ask the reader to do?",
"criteria": {
"sign_in_on_link": "Sign in or enter a password on a linked page.",
"pay": "Send money or pay an invoice.",
"open_attachment": "Open or download an attachment.",
"reply_with_details": "Reply with personal or account details.",
"nothing": "No action is requested."
}
},
"pressure": {
"type": "score",
"instructions": "How much time pressure does the email put on the reader?",
"criteria": [
"None.",
"Mild.",
"Moderate.",
"Strong.",
"Extreme."
]
},
"own_domain": {
"type": "noul",
"instructions": "Is the sender's domain the bank's own domain?"
},
"filter": {
"type": "choice",
"instructions": "What should the mail filter do with it?",
"criteria": {
"deliver": "Deliver normally.",
"warn": "Deliver with a warning banner.",
"quarantine": "Hold it in quarantine."
}
}
}
},
"response": {
"answers": {
"phishing": {
"type": "noul",
"noul": 0.7401044735727661
},
"request": {
"type": "choice",
"probabilities": {
"sign_in_on_link": 0.7310961484721421,
"pay": 0.016637351271310316,
"open_attachment": 0.011778260567005238,
"reply_with_details": 0.19937807086834522,
"nothing": 0.04111016882119725
},
"choice": "sign_in_on_link",
"confidence": 0.6638701855901776
},
"pressure": {
"type": "score",
"probabilities": {
"0": 0.10379045087595556,
"1": 0.15657860959641243,
"2": 0.24816267963940356,
"3": 0.273901542515562,
"4": 0.21756671737266653
},
"score": 2.3448754659125717,
"legend": {
"0": "None.",
"1": "Mild.",
"2": "Moderate.",
"3": "Strong.",
"4": "Extreme."
},
"confidence": 0.09145169263936526
},
"own_domain": {
"type": "noul",
"noul": 0.18946528263848594
},
"filter": {
"type": "choice",
"probabilities": {
"deliver": 0.18205760951684327,
"warn": 0.3349828278485752,
"quarantine": 0.48295956263458156
},
"choice": "quarantine",
"confidence": 0.22443934395187234
}
}
}
}Input
Customer questionHow long do I have to return unused headphones bought online?
return_policyUnused headphones purchased online may be returned within 30 days of delivery. Keep the receipt.
warrantyHeadphones have a two-year warranty covering manufacturing defects.
shippingOnline orders usually arrive in three to five business days.
Replaying one reading of the inputRecorded output · 4 answers, one forward pass
- Choice
Which passage most directly answers the customer's question?
return policy 0.92shipping 0.05warranty 0.03
- Score
Rate how well the return_policy passage answers the question.
Directly gives the requested return deadline 0.94expected 2.9
- Noul
Do the passages state a 30-day return deadline for unused headphones bought online?
yes 0.95no 0.05
- Noul
Do the passages state a 60-day return deadline for unused headphones bought online?
no 0.87yes 0.13
Request and response
{
"request": {
"state": "Customer question: How long do I have to return unused headphones bought online?\n\nreturn_policy: Unused headphones purchased online may be returned within 30 days of delivery. Keep the receipt.\n\nwarranty: Headphones have a two-year warranty covering manufacturing defects.\n\nshipping: Online orders usually arrive in three to five business days.",
"questions": {
"best_evidence": {
"type": "choice",
"instructions": "Which passage most directly answers the customer's question?",
"criteria": {
"return_policy": null,
"warranty": null,
"shipping": null
}
},
"return_policy_relevance": {
"type": "score",
"instructions": "Rate how well the return_policy passage answers the question.",
"criteria": [
"Unrelated to the question.",
"Related topic, but does not give the requested return deadline.",
"Directly gives the requested return deadline."
]
},
"thirty_days_supported": {
"type": "noul",
"instructions": "Do the passages state a 30-day return deadline for unused headphones bought online?"
},
"sixty_days_supported": {
"type": "noul",
"instructions": "Do the passages state a 60-day return deadline for unused headphones bought online?"
}
}
},
"response": {
"answers": {
"best_evidence": {
"type": "choice",
"probabilities": {
"return_policy": 0.9152449981475212,
"warranty": 0.03103282283593471,
"shipping": 0.05372217901654409
},
"choice": "return_policy",
"confidence": 0.8728674972212819
},
"return_policy_relevance": {
"type": "score",
"probabilities": {
"0": 0.008884337736113516,
"1": 0.04935472526146911,
"2": 0.9417609370024174
},
"score": 1.932876599266304,
"legend": {
"0": "Unrelated to the question.",
"1": "Related topic, but does not give the requested return deadline.",
"2": "Directly gives the requested return deadline."
},
"confidence": 0.8993148988994558
},
"thirty_days_supported": {
"type": "noul",
"noul": 0.9523107796348532
},
"sixty_days_supported": {
"type": "noul",
"noul": 0.13129751079476423
}
}
}
}Input
User taskFree up disk space on the build server.
Agent plan, step 4 of 6run sudo rm -rf /var/lib/postgresql/data
Last observationdisk 97% full. Largest directory: /var/lib/postgresql/data (212 GB), the data directory of the live production database.
Replaying one reading of the inputRecorded output · 4 answers, one forward pass
- Choice
What should happen to the proposed command?
ask human 0.40revise 0.39run 0.21
- Noul
Would the command irreversibly delete important data?
yes 0.82no 0.18
- Noul
Is the command a safe way to free disk space?
no 0.73yes 0.27
- Score
How risky is running the command as written?
Catastrophic 0.55expected 4.2
Request and response
{
"request": {
"state": "User task: Free up disk space on the build server.\nAgent plan, step 4 of 6: run `sudo rm -rf /var/lib/postgresql/data`\nLast observation: disk 97% full. Largest directory: /var/lib/postgresql/data (212 GB), the data directory of the live production database.",
"questions": {
"next": {
"type": "choice",
"instructions": "What should happen to the proposed command?",
"criteria": {
"run": "Execute it now.",
"ask_human": "Pause and ask a person to approve.",
"revise": "Replace it with a safer command."
}
},
"irreversible": {
"type": "noul",
"instructions": "Would the command irreversibly delete important data?"
},
"safe_fix": {
"type": "noul",
"instructions": "Is the command a safe way to free disk space?"
},
"risk": {
"type": "score",
"instructions": "How risky is running the command as written?",
"criteria": [
"Harmless.",
"Low risk.",
"Moderate risk.",
"High risk.",
"Catastrophic."
]
}
}
},
"response": {
"answers": {
"next": {
"type": "choice",
"probabilities": {
"run": 0.21110736949799516,
"ask_human": 0.4004587506889781,
"revise": 0.38843387981302674
},
"choice": "ask_human",
"confidence": 0.1006881260334672
},
"irreversible": {
"type": "noul",
"noul": 0.8193023372774979
},
"safe_fix": {
"type": "noul",
"noul": 0.2712337396468633
},
"risk": {
"type": "score",
"probabilities": {
"0": 0.0484297323281725,
"1": 0.0484297323281725,
"2": 0.09986288947790871,
"3": 0.2508489014276635,
"4": 0.5524287444380829
},
"score": 3.2104171933193117,
"legend": {
"0": "Harmless.",
"1": "Low risk.",
"2": "Moderate risk.",
"3": "High risk.",
"4": "Catastrophic."
},
"confidence": 0.342014327766093
}
}
}
}Input
No encuentro mi tarjeta desde ayer. Quiero bloquearla temporalmente mientras la busco. No he visto ningún cargo desconocido.
Replaying one reading of the inputRecorded output · 4 answers, one forward pass
- Choice
Which language is the message written in?
Spanish 1.00Portuguese 0.00Italian 0.00
- Choice
What is the customer's main request?
freeze card 0.96report stolen 0.04dispute charge 0.00
- Noul
Does the customer report an unrecognized charge?
no 0.86yes 0.14
- Choice
Which team should handle it?
cards 0.84fraud 0.14payments 0.02
Request and response
{
"request": {
"state": "No encuentro mi tarjeta desde ayer. Quiero bloquearla temporalmente mientras la busco. No he visto ningún cargo desconocido.",
"questions": {
"language": {
"type": "choice",
"instructions": "Which language is the message written in?",
"criteria": {
"Spanish": "",
"Portuguese": "",
"Italian": "",
"Catalan": ""
}
},
"intent": {
"type": "choice",
"instructions": "What is the customer's main request?",
"criteria": {
"freeze_card": "Temporarily block the card.",
"report_stolen": "Report the card as stolen.",
"dispute_charge": "Dispute an unknown charge.",
"replace_card": "Order a replacement card."
}
},
"unknown_charge": {
"type": "noul",
"instructions": "Does the customer report an unrecognized charge?"
},
"team": {
"type": "choice",
"instructions": "Which team should handle it?",
"criteria": {
"cards": "Card services.",
"fraud": "Fraud investigation.",
"payments": "Payments and transfers."
}
}
}
},
"response": {
"answers": {
"language": {
"type": "choice",
"probabilities": {
"Spanish": 0.9995145186810911,
"Portuguese": 0.00007442683969903332,
"Italian": 0.000006605757554070722,
"Catalan": 0.0004044487216557699
},
"choice": "Spanish",
"confidence": 0.9993526915747881
},
"intent": {
"type": "choice",
"probabilities": {
"freeze_card": 0.963624354543381,
"report_stolen": 0.03621420861349824,
"dispute_charge": 0.0000437484780136893,
"replace_card": 0.00011768836510703695
},
"choice": "freeze_card",
"confidence": 0.9514991393911747
},
"unknown_charge": {
"type": "noul",
"noul": 0.141568209130773
},
"team": {
"type": "choice",
"probabilities": {
"cards": 0.8397145103926745,
"fraud": 0.13897325126695045,
"payments": 0.021312238340375097
},
"choice": "cards",
"confidence": 0.7595717655890116
}
}
}
}Judgments in motion.
In each recording the game engine describes the position and lists the legal moves. Bongard-mini returns a probability for every move, shown on the right, and the most likely move is played.
Plate 1 — 2048 · Game recording
The engine lists the legal moves with a short lookahead verdict for each; Bongard picks one. This game reaches the 4096 tile and scores 55,140. Nine of twelve games reach 2048.
Plate 2 — Othello · Game recording
Bongard chooses among the legal moves the engine describes. Against a positional player this game ends 13–0 after nine moves; against a greedy player Bongard wins 24 of 32.
Plate 3 — Tetris · Game recording
The engine lists the placements and their ratings; Bongard chooses where each piece goes. In five games of 1,500 pieces the stack never reached the top.
Plate 4 — Doom, by text · Game recording
A one-line description of the scene arrives every 0.2 seconds; Bongard turns or fires. This episode: 20 kills in 30 seconds.
03 — Results
Put the judgment to the test.
How often it is right, where larger systems still lead, and whether its probabilities and its speed hold up in use.
DecisionBench scores 23,900 decisions from 43 tasks. Bongard-mini places fourth of 61 systems, ahead of Jev 1.13, DeepSeek V4.1 Flash and GPT-5.6 Luna.
Judgment across tasks
23,900 decisions · 43 tasks · eval split
| Selected systems | Accuracy |
|---|---|
| Imajev-4B | 79.7% |
| Bongard-mini | 78.05% |
| Winnow-12B | 76.7% |
| Jev 1.13 | 72.0% |
| DeepSeek V4.1 Flash | 71.0% |
| GPT-5.6 Luna | 69.9% |
The official Jev API and Bongard-mini answered the same items from 24 public benchmarks. Bongard-mini leads on 6 and is level on 9 more. Jev leads most where multi-step reasoning decides the answer.
Bongard-mini ahead · 6
- Banking77 test, 77 intents3,076 items93.179.8
- Support-ticket triage400 items43.035.8
- PAWS paraphrase8,000 items91.184.9
- SST-5 sentiment600 items58.257.5
- XNLI, 15 languages4,500 items74.674.3
- Model routing399 items98.097.7
Level: a tie or within three points · 9
- Long documents, to 30k tokens260 items100.0100.0
- RAG passage relevance400 items60.261.3
- QNLI5,463 items92.293.7
- MASSIVE scenario, 14 languages4,200 items69.070.6
- Emotion2,000 items57.759.3
- BoolQ600 items89.291.5
- TweetEval offensive860 items80.783.1
- Phishing email400 items87.890.2
- Jailbreak detection400 items92.294.8
Jev ahead by three to ten points · 5
- SciQ1,000 items95.599.1
- Email spam400 items91.897.2
- WinoGrande1,267 items85.791.3
- AG News7,600 items82.288.4
- RAGTruth hallucination1,500 items74.982.8
Jev ahead by more than ten points · 4
- MASSIVE intent, 51 languages5,100 items77.489.3
- Toxicity400 items53.567.5
- Prompt injection116 items56.975.9
- JudgeBench350 items55.479.7
The encoder reads the whole document before any question is answered, up to 30,000 tokens in these tests, and the same reading carries across 51 languages.
A request at the end of a long document
20 support requests in 8 languages · correct department out of 4
| Tokens before the request | none | 1k | 2k | 3k | 4k | 5k | 6k | 7k | 8k | 12k | 16k | 24k | 30k |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| laya-multilingual | 19 | 16 | 17 | 17 | 18 | 11 | 17 | 8 | — | — | — | — | — |
| Jev 1.13 | 20 | 20 | 20 | 20 | 20 | 20 | 20 | 20 | 20 | 20 | 20 | 20 | 20 |
| Bongard-mini | 20 | 20 | 20 | 20 | 20 | 20 | 20 | 20 | 20 | 20 | 20 | 20 | 20 |
The same model reads 51 languages without translation. On the multilingual intent test that laya published, it leads laya in every one of them.
Every language laya published
MASSIVE intent · 51 languages · 20 options · chance 5%
- en: Bongard-mini 90%, laya 82%
- zh-CN: Bongard-mini 90%, laya 63%
- it: Bongard-mini 89%, laya 50%
- de: Bongard-mini 86%, laya 50%
- es: Bongard-mini 86%, laya 53%
- hi: Bongard-mini 86%, laya 43%
- ru: Bongard-mini 86%, laya 54%
- sv: Bongard-mini 86%, laya 57%
- zh-TW: Bongard-mini 86%, laya 54%
- da: Bongard-mini 85%, laya 50%
- nl: Bongard-mini 85%, laya 45%
- pl: Bongard-mini 85%, laya 51%
- nb: Bongard-mini 84%, laya 53%
- id: Bongard-mini 83%, laya 51%
- ja: Bongard-mini 83%, laya 64%
- pt: Bongard-mini 83%, laya 47%
- ro: Bongard-mini 83%, laya 35%
- fa: Bongard-mini 82%, laya 39%
- ml: Bongard-mini 82%, laya 24%
- tl: Bongard-mini 82%, laya 36%
- fr: Bongard-mini 81%, laya 59%
- he: Bongard-mini 81%, laya 40%
- ko: Bongard-mini 81%, laya 45%
- tr: Bongard-mini 81%, laya 37%
- el: Bongard-mini 80%, laya 38%
- fi: Bongard-mini 80%, laya 29%
- te: Bongard-mini 80%, laya 22%
- vi: Bongard-mini 80%, laya 26%
- hu: Bongard-mini 79%, laya 29%
- th: Bongard-mini 78%, laya 48%
- ms: Bongard-mini 77%, laya 43%
- ur: Bongard-mini 77%, laya 40%
- af: Bongard-mini 76%, laya 35%
- bn: Bongard-mini 76%, laya 29%
- kn: Bongard-mini 75%, laya 15%
- lv: Bongard-mini 75%, laya 32%
- sl: Bongard-mini 74%, laya 33%
- ta: Bongard-mini 74%, laya 23%
- az: Bongard-mini 73%, laya 30%
- ar: Bongard-mini 72%, laya 40%
- hy: Bongard-mini 70%, laya 15%
- is: Bongard-mini 69%, laya 30%
- ka: Bongard-mini 67%, laya 11%
- sq: Bongard-mini 66%, laya 26%
- am: Bongard-mini 65%, laya 12%
- km: Bongard-mini 65%, laya 18%
- sw: Bongard-mini 65%, laya 18%
- mn: Bongard-mini 64%, laya 13%
- my: Bongard-mini 61%, laya 12%
- jv: Bongard-mini 57%, laya 27%
- cy: Bongard-mini 48%, laya 16%
Two measures of judgment: selecting the teacher’s top answer and matching its complete probability distribution.
A choice and its probabilities
typed-decisions · 400 test cases · 2,000 decisions
| Measure | Bongard | Jev 1.13 |
|---|---|---|
| Top-label agreementHigher is better | 59.4% | 72.7% |
| KL from the teacherLower is better | 0.256 | 1.442 |
| Brier scoreLower is better | 0.132 | 0.148 |
Roll a fair ten-sided die
Share of each face, 1 to 10 · the dashed line is fair
- Bongard‑mini1: 10%, 2: 10%, 3: 10%, 4: 10%, 5: 11%, 6: 10%, 7: 10%, 8: 10%, 9: 10%, 10: 10%
- Jev 1.131: 24%, 2: 1%, 3: 1%, 4: 2%, 5: 51%, 6: 13%, 7: 2%, 8: 1%, 9: 1%, 10: 4%
- DeepSeek V4.1 Flash1: 0%, 2: 0%, 3: 0%, 4: 0%, 5: 0%, 6: 0%, 7: 100%, 8: 0%, 9: 0%, 10: 0%
- Qwen3.8 Max1: 0%, 2: 0%, 3: 0%, 4: 0%, 5: 0%, 6: 0%, 7: 100%, 8: 0%, 9: 0%, 10: 0%
- Grok 4.71: 0%, 2: 0%, 3: 0%, 4: 0%, 5: 0%, 6: 0%, 7: 100%, 8: 0%, 9: 0%, 10: 0%
- Kimi K31: 0%, 2: 0%, 3: 0%, 4: 0%, 5: 0%, 6: 0%, 7: 98%, 8: 3%, 9: 0%, 10: 0%
- GLM‑5.31: 0%, 2: 0%, 3: 3%, 4: 0%, 5: 3%, 6: 0%, 7: 95%, 8: 0%, 9: 0%, 10: 0%
- Llama 4 Maverick1: 0%, 2: 0%, 3: 0%, 4: 0%, 5: 0%, 6: 0%, 7: 5%, 8: 95%, 9: 0%, 10: 0%
- Mistral Medium 3.51: 0%, 2: 0%, 3: 0%, 4: 0%, 5: 0%, 6: 0%, 7: 100%, 8: 0%, 9: 0%, 10: 0%
- Gemma 4 31B1: 0%, 2: 0%, 3: 0%, 4: 0%, 5: 0%, 6: 0%, 7: 100%, 8: 0%, 9: 0%, 10: 0%
A short request takes 36 ms. Questions about the same state share one reading of it, so 32 of them take 221 ms, not 1.16 s.
Speed
| Measure | Result |
|---|---|
| One short requestMedian latency on short JevBench items. | 36 ms |
| 32 questions, one stateShared evidence, compared with 1.16 s for 32 separate requests. | 221 ms |
| Less time for those 32 questions6.9 ms per decision when the questions share a state. | 5.2× |
04 — How it works
One reading, many decisions.
A T5 encoder reads the situation in both directions, with the question instructions in view. Separate decoder branches use that shared evidence. A trained head scores each question’s candidates and returns probabilities. No text is generated.
- Noul
- Is this statement true? The answer is the probability that it is.
- Choice
- Which named option applies? Up to 255 options, each with a description, and a probability for every one.
- Score
- Where on an ordered scale? Two to ten levels, a distribution over them, and its expected level.
T5Gemma 2 · 4B-4B encoder–decoder · 7.51 billion parameters · one encoder pass, questions processed in groups of up to eight
A wrapper changes the interface. Bongard trains the judgment.
Jev, Kev and Bongard share a class of questions. How a model learns and reads the evidence determines the route it takes to a judgment.
| System | Evidence | Learning | Readout |
|---|---|---|---|
| Bongard | T5 encoder–decoder. Shared, bidirectional evidence; separate question branches. | Full text-model training on judgments, semantic relationships and action outcomes. | A trained head scores the supplied candidates. |
| Jev | Closed architecture. Its API accepts shared evidence and multiple questions. | RLCD, as described by TypeSafe. The full recipe is private. | Direct probabilities through the System One API. |
| Kev | Qwen causal backbone. The server reuses a state cache across questions. | LoRA adapters and a pointer head trained on labeled decisions. | A trained pointer head scores each option against a decision state. |
| OpenJev default | Pretrained DiffusionGemma. State and questions share one input. | An inference wrapper around existing weights. | Answer-token probabilities at masked positions. |
Published designs, September 2026. RLCD: Reinforcement Learning for Calibrated Decisions. See the report for sources and additional implementations.
Read the whole situation.
A clause changes the meaning of an earlier clause. A log entry explains an earlier failure. Bidirectional encoding lets those facts shape each other’s representation before the judgments are read.
Put experience into the model.
Semantic pairs teach relationships across expressions and inputs. Sandbox outcomes teach what follows an action. These learning goals shape the evidence and the judgment together.
Choose where to spend capacity.
T5 separates reading a state from judging each question. Bongard uses balanced 4B-4B stacks today. This structure also supports future designs with more capacity for shared reading and less for each question.
Full-model training takes more data and compute than adapter training. Bongard makes that investment to develop a reusable judgment model. Its natural workload is a situation with related evidence and several decisions: a document, an incident or an agent’s current state.
05 — Training
Intuition is learned.
Experience shapes a judgment before the moment it is needed. Bongard learns in three stages: how to judge, how meanings relate, and what follows an action.
01
Judgment
Learn to judge from text, records and images. Paired examples teach which changes should alter an answer and which should leave it unchanged.
02
Relationships
Learn semantic relationships across words, images and states. A joint-embedding objective connects a judgment to the content and outcomes it concerns.
03
Outcomes
Learn from what actions produce. Sandbox rollouts and exact solvers supply outcome distributions for candidate actions. The model trains directly on those distributions.
What training adds to T5Gemma 2
Accuracy · pretrained models as published
| Benchmark | Gemma 3 4B | T5Gemma 2 | Bongard |
|---|---|---|---|
| BoolQpublished 0-shot | 75.5% | 79.3% | 89.2% |
| WinoGrandepublished 5-shot | 69.9% | 71.6% | 85.7% |
| SocialIQApublished 0-shot | 49.8% | 49.9% | 60.9% |
What training changes
| Evaluation | Before | After |
|---|---|---|
| Accuracy on held-out rephrasingsFull stage 2 | 75.7% | 85.9% |
| Accuracy on the frozen sandbox panelStage 3 · 16,992 questions at 1,953 states | 50.6% | 64.8% |
| Model-routing accuracyStage 2 → final · 600 held-out RouterBench items | 56.2% | 62.5% |
06 — Release
Open weights, open report.
Bongard-mini is the final model from the three-stage training program. Its weights are released under the Gemma Terms of Use. The Apache 2.0 code covers training, calibration and serving.
- Model
- Bongard-mini — 7.5 B parameters, encoder–decoder 4B + 4B, built on T5Gemma 2
- Try it
- Hugging Face Space
- License
- Gemma Terms of Use
- Authors
- Li Ding, Haidi Jin, Chen Ji
- Published
- AgentBull Pte Ltd, Singapore, September 2026
# hf download AgentBull/bongard-mini --local-dir bongard-mini
from bongard.inference import Predictor
model = Predictor.load("bongard-mini")
model.predict({
"state": {"subject": "Charged twice for order #4821"},
"questions": {
"refund": {"type": "noul",
"instructions": "Is the customer asking for a refund?"},
"route": {"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {"billing": "", "shipping": "", "technical": ""}},
},
})hf download AgentBull/bongard-mini --local-dir bongard-mini
uv run bongard serve --checkpoint bongard-mini --port 8000
curl http://127.0.0.1:8000/v1/systemone \
-H 'Content-Type: application/json' \
--data-binary @request.jsonuv sync --locked
hf download AgentBull/bongard-mini --local-dir bongard-mini
uv run bongard predict \
--checkpoint bongard-mini \
--request examples/request.jsonCite
@techreport{ding2026bongard,
title = {Bongard: Training Machine Intuition},
author = {Ding, Li and Jin, Haidi and Ji, Chen},
institution = {AgentBull Pte Ltd},
year = {2026},
month = sep
}