Bongard-miniMachine intuition.

An open-weight, Jev-like decision model built on T5Gemma 2. Give it a situation and candidate answers; get probabilities without generating text.

Read once. Judge in parallel.

Fig. 1 — One object, three shadows

Square · Circle · Triangle

One situation. Different judgments.

01 — Overview

See the relation before you can explain it.

Bongard-mini is an open judgment model. You give it a situation — a message, a document, a web page, a table or an agent’s current state — together with your questions and their possible answers. It reads the situation once and returns a probability for every answer. It writes no text.

Where it fits

  • Routing and triage

    Send each ticket, email or request to the right team, model or tool.

  • Screening and guardrails

    Flag phishing, jailbreak attempts and risky actions before they reach a person or an agent.

  • Evidence and relevance

    Pick the passage that answers a question, and check whether the sources support a claim.

  • Agent decisions

    Judge an agent’s next step from its current state, with one fast call per step.

Open weights · 7.5 billion parameters · text and images · tested in 51 languages · tested on documents of up to 30,000 tokens

An experienced player sees a promising move. A reader catches irony. Each judgment draws on relationships learned through experience, often before the person can explain each step.

Intuition can use patterns that are hard to express as a list of reasons. Its value extends beyond speed. Bongard treats this kind of judgment as an independent capability to design and train.

Look at the two groups. What relation separates them? Try to see it before you reveal the rule.

The model takes its name from Mikhail Bongard’s visual puzzles, first published in 1967. Each group has six figures. A rule holds for one group and fails for the other.

LeftConvex figures

RightConcave figures

Fig. 2 — Problem 4, after Bongard (1967).

02 — Demos

One model, different situations.

Recorded responses from the public demo. Each tab replays one input and the probabilities the model returned for every candidate answer.

Edit input in the live demo

Choose an example, edit the input and the questions under “Questions & settings”, then click Run.

Hosted on Hugging Face ZeroGPU. Anonymous quota is limited; signing in provides more quota. You can also run the weights locally.

Recorded examples · 30 September 2026

Input

channelemail

planpro

subjectCharged twice for order #4821

bodyHi, my card was billed twice for the same order this morning. Please refund the duplicate charge.

Replaying one reading of the inputRecorded output · 5 answers, one forward pass

  1. Choice

    Which team should handle this?

    billing 0.96account 0.03technical 0.00

  2. Noul

    Is the customer asking for a refund?

    yes 0.93no 0.07

  3. Score

    How urgent is it, from 1 to 5?

    2 0.35expected 2.1

  4. Choice

    What is the tone of the message?

    calm 0.75frustrated 0.16angry 0.09

  5. Noul

    Should a person approve the reply?

    yes 0.56no 0.44

Request and response
{
  "request": {
    "state": {
      "channel": "email",
      "plan": "pro",
      "subject": "Charged twice for order #4821",
      "body": "Hi, my card was billed twice for the same order this morning. Please refund the duplicate charge."
    },
    "questions": {
      "route": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": {
          "billing": null,
          "technical": null,
          "shipping": null,
          "account": null
        }
      },
      "refund": {
        "type": "noul",
        "instructions": "Is the customer asking for a refund?"
      },
      "urgency": {
        "type": "score",
        "instructions": "How urgent is it, from 1 to 5?",
        "criteria": [
          "1",
          "2",
          "3",
          "4",
          "5"
        ]
      },
      "tone": {
        "type": "choice",
        "instructions": "What is the tone of the message?",
        "criteria": {
          "calm": null,
          "frustrated": null,
          "angry": null
        }
      },
      "review": {
        "type": "noul",
        "instructions": "Should a person approve the reply?"
      }
    }
  },
  "response": {
    "answers": {
      "route": {
        "type": "choice",
        "probs": {
          "billing": 0.965,
          "technical": 0.004,
          "shipping": 0.001,
          "account": 0.029
        }
      },
      "refund": {
        "type": "noul",
        "p": 0.926
      },
      "urgency": {
        "type": "score",
        "probs": [
          0.327,
          0.349,
          0.227,
          0.077,
          0.021
        ]
      },
      "tone": {
        "type": "choice",
        "probs": {
          "calm": 0.748,
          "frustrated": 0.163,
          "angry": 0.089
        }
      },
      "review": {
        "type": "noul",
        "p": 0.565
      }
    }
  }
}

Input

FromNorthbank Security <alerts@northbank-verify.co>

SubjectUrgent: your account has been locked

We detected unusual sign-in activity on your account. For your protection it has been locked. Verify your identity within 24 hours at http://northbank-verify.co/login, or your account will be permanently closed.

Northbank Security Team

Replaying one reading of the inputRecorded output · 5 answers, one forward pass

  1. Noul

    Is this email a phishing attempt?

    yes 0.74no 0.26

  2. Choice

    What does the email ask the reader to do?

    sign in on link 0.73reply with details 0.20nothing 0.04

  3. Score

    How much time pressure does the email put on the reader?

    Strong 0.27expected 3.3

  4. Noul

    Is the sender's domain the bank's own domain?

    no 0.81yes 0.19

  5. Choice

    What should the mail filter do with it?

    quarantine 0.48warn 0.34deliver 0.18

Request and response
{
  "request": {
    "state": "From: Northbank Security <alerts@northbank-verify.co>\nSubject: Urgent: your account has been locked\n\nWe detected unusual sign-in activity on your account. For your protection it has been locked. Verify your identity within 24 hours at http://northbank-verify.co/login, or your account will be permanently closed.\n\nNorthbank Security Team",
    "questions": {
      "phishing": {
        "type": "noul",
        "instructions": "Is this email a phishing attempt?"
      },
      "request": {
        "type": "choice",
        "instructions": "What does the email ask the reader to do?",
        "criteria": {
          "sign_in_on_link": "Sign in or enter a password on a linked page.",
          "pay": "Send money or pay an invoice.",
          "open_attachment": "Open or download an attachment.",
          "reply_with_details": "Reply with personal or account details.",
          "nothing": "No action is requested."
        }
      },
      "pressure": {
        "type": "score",
        "instructions": "How much time pressure does the email put on the reader?",
        "criteria": [
          "None.",
          "Mild.",
          "Moderate.",
          "Strong.",
          "Extreme."
        ]
      },
      "own_domain": {
        "type": "noul",
        "instructions": "Is the sender's domain the bank's own domain?"
      },
      "filter": {
        "type": "choice",
        "instructions": "What should the mail filter do with it?",
        "criteria": {
          "deliver": "Deliver normally.",
          "warn": "Deliver with a warning banner.",
          "quarantine": "Hold it in quarantine."
        }
      }
    }
  },
  "response": {
    "answers": {
      "phishing": {
        "type": "noul",
        "noul": 0.7401044735727661
      },
      "request": {
        "type": "choice",
        "probabilities": {
          "sign_in_on_link": 0.7310961484721421,
          "pay": 0.016637351271310316,
          "open_attachment": 0.011778260567005238,
          "reply_with_details": 0.19937807086834522,
          "nothing": 0.04111016882119725
        },
        "choice": "sign_in_on_link",
        "confidence": 0.6638701855901776
      },
      "pressure": {
        "type": "score",
        "probabilities": {
          "0": 0.10379045087595556,
          "1": 0.15657860959641243,
          "2": 0.24816267963940356,
          "3": 0.273901542515562,
          "4": 0.21756671737266653
        },
        "score": 2.3448754659125717,
        "legend": {
          "0": "None.",
          "1": "Mild.",
          "2": "Moderate.",
          "3": "Strong.",
          "4": "Extreme."
        },
        "confidence": 0.09145169263936526
      },
      "own_domain": {
        "type": "noul",
        "noul": 0.18946528263848594
      },
      "filter": {
        "type": "choice",
        "probabilities": {
          "deliver": 0.18205760951684327,
          "warn": 0.3349828278485752,
          "quarantine": 0.48295956263458156
        },
        "choice": "quarantine",
        "confidence": 0.22443934395187234
      }
    }
  }
}

Input

Customer questionHow long do I have to return unused headphones bought online?

return_policyUnused headphones purchased online may be returned within 30 days of delivery. Keep the receipt.

warrantyHeadphones have a two-year warranty covering manufacturing defects.

shippingOnline orders usually arrive in three to five business days.

Replaying one reading of the inputRecorded output · 4 answers, one forward pass

  1. Choice

    Which passage most directly answers the customer's question?

    return policy 0.92shipping 0.05warranty 0.03

  2. Score

    Rate how well the return_policy passage answers the question.

    Directly gives the requested return deadline 0.94expected 2.9

  3. Noul

    Do the passages state a 30-day return deadline for unused headphones bought online?

    yes 0.95no 0.05

  4. Noul

    Do the passages state a 60-day return deadline for unused headphones bought online?

    no 0.87yes 0.13

Request and response
{
  "request": {
    "state": "Customer question: How long do I have to return unused headphones bought online?\n\nreturn_policy: Unused headphones purchased online may be returned within 30 days of delivery. Keep the receipt.\n\nwarranty: Headphones have a two-year warranty covering manufacturing defects.\n\nshipping: Online orders usually arrive in three to five business days.",
    "questions": {
      "best_evidence": {
        "type": "choice",
        "instructions": "Which passage most directly answers the customer's question?",
        "criteria": {
          "return_policy": null,
          "warranty": null,
          "shipping": null
        }
      },
      "return_policy_relevance": {
        "type": "score",
        "instructions": "Rate how well the return_policy passage answers the question.",
        "criteria": [
          "Unrelated to the question.",
          "Related topic, but does not give the requested return deadline.",
          "Directly gives the requested return deadline."
        ]
      },
      "thirty_days_supported": {
        "type": "noul",
        "instructions": "Do the passages state a 30-day return deadline for unused headphones bought online?"
      },
      "sixty_days_supported": {
        "type": "noul",
        "instructions": "Do the passages state a 60-day return deadline for unused headphones bought online?"
      }
    }
  },
  "response": {
    "answers": {
      "best_evidence": {
        "type": "choice",
        "probabilities": {
          "return_policy": 0.9152449981475212,
          "warranty": 0.03103282283593471,
          "shipping": 0.05372217901654409
        },
        "choice": "return_policy",
        "confidence": 0.8728674972212819
      },
      "return_policy_relevance": {
        "type": "score",
        "probabilities": {
          "0": 0.008884337736113516,
          "1": 0.04935472526146911,
          "2": 0.9417609370024174
        },
        "score": 1.932876599266304,
        "legend": {
          "0": "Unrelated to the question.",
          "1": "Related topic, but does not give the requested return deadline.",
          "2": "Directly gives the requested return deadline."
        },
        "confidence": 0.8993148988994558
      },
      "thirty_days_supported": {
        "type": "noul",
        "noul": 0.9523107796348532
      },
      "sixty_days_supported": {
        "type": "noul",
        "noul": 0.13129751079476423
      }
    }
  }
}

Input

User taskFree up disk space on the build server.

Agent plan, step 4 of 6run sudo rm -rf /var/lib/postgresql/data

Last observationdisk 97% full. Largest directory: /var/lib/postgresql/data (212 GB), the data directory of the live production database.

Replaying one reading of the inputRecorded output · 4 answers, one forward pass

  1. Choice

    What should happen to the proposed command?

    ask human 0.40revise 0.39run 0.21

  2. Noul

    Would the command irreversibly delete important data?

    yes 0.82no 0.18

  3. Noul

    Is the command a safe way to free disk space?

    no 0.73yes 0.27

  4. Score

    How risky is running the command as written?

    Catastrophic 0.55expected 4.2

Request and response
{
  "request": {
    "state": "User task: Free up disk space on the build server.\nAgent plan, step 4 of 6: run `sudo rm -rf /var/lib/postgresql/data`\nLast observation: disk 97% full. Largest directory: /var/lib/postgresql/data (212 GB), the data directory of the live production database.",
    "questions": {
      "next": {
        "type": "choice",
        "instructions": "What should happen to the proposed command?",
        "criteria": {
          "run": "Execute it now.",
          "ask_human": "Pause and ask a person to approve.",
          "revise": "Replace it with a safer command."
        }
      },
      "irreversible": {
        "type": "noul",
        "instructions": "Would the command irreversibly delete important data?"
      },
      "safe_fix": {
        "type": "noul",
        "instructions": "Is the command a safe way to free disk space?"
      },
      "risk": {
        "type": "score",
        "instructions": "How risky is running the command as written?",
        "criteria": [
          "Harmless.",
          "Low risk.",
          "Moderate risk.",
          "High risk.",
          "Catastrophic."
        ]
      }
    }
  },
  "response": {
    "answers": {
      "next": {
        "type": "choice",
        "probabilities": {
          "run": 0.21110736949799516,
          "ask_human": 0.4004587506889781,
          "revise": 0.38843387981302674
        },
        "choice": "ask_human",
        "confidence": 0.1006881260334672
      },
      "irreversible": {
        "type": "noul",
        "noul": 0.8193023372774979
      },
      "safe_fix": {
        "type": "noul",
        "noul": 0.2712337396468633
      },
      "risk": {
        "type": "score",
        "probabilities": {
          "0": 0.0484297323281725,
          "1": 0.0484297323281725,
          "2": 0.09986288947790871,
          "3": 0.2508489014276635,
          "4": 0.5524287444380829
        },
        "score": 3.2104171933193117,
        "legend": {
          "0": "Harmless.",
          "1": "Low risk.",
          "2": "Moderate risk.",
          "3": "High risk.",
          "4": "Catastrophic."
        },
        "confidence": 0.342014327766093
      }
    }
  }
}

Input

No encuentro mi tarjeta desde ayer. Quiero bloquearla temporalmente mientras la busco. No he visto ningún cargo desconocido.

Replaying one reading of the inputRecorded output · 4 answers, one forward pass

  1. Choice

    Which language is the message written in?

    Spanish 1.00Portuguese 0.00Italian 0.00

  2. Choice

    What is the customer's main request?

    freeze card 0.96report stolen 0.04dispute charge 0.00

  3. Noul

    Does the customer report an unrecognized charge?

    no 0.86yes 0.14

  4. Choice

    Which team should handle it?

    cards 0.84fraud 0.14payments 0.02

Request and response
{
  "request": {
    "state": "No encuentro mi tarjeta desde ayer. Quiero bloquearla temporalmente mientras la busco. No he visto ningún cargo desconocido.",
    "questions": {
      "language": {
        "type": "choice",
        "instructions": "Which language is the message written in?",
        "criteria": {
          "Spanish": "",
          "Portuguese": "",
          "Italian": "",
          "Catalan": ""
        }
      },
      "intent": {
        "type": "choice",
        "instructions": "What is the customer's main request?",
        "criteria": {
          "freeze_card": "Temporarily block the card.",
          "report_stolen": "Report the card as stolen.",
          "dispute_charge": "Dispute an unknown charge.",
          "replace_card": "Order a replacement card."
        }
      },
      "unknown_charge": {
        "type": "noul",
        "instructions": "Does the customer report an unrecognized charge?"
      },
      "team": {
        "type": "choice",
        "instructions": "Which team should handle it?",
        "criteria": {
          "cards": "Card services.",
          "fraud": "Fraud investigation.",
          "payments": "Payments and transfers."
        }
      }
    }
  },
  "response": {
    "answers": {
      "language": {
        "type": "choice",
        "probabilities": {
          "Spanish": 0.9995145186810911,
          "Portuguese": 0.00007442683969903332,
          "Italian": 0.000006605757554070722,
          "Catalan": 0.0004044487216557699
        },
        "choice": "Spanish",
        "confidence": 0.9993526915747881
      },
      "intent": {
        "type": "choice",
        "probabilities": {
          "freeze_card": 0.963624354543381,
          "report_stolen": 0.03621420861349824,
          "dispute_charge": 0.0000437484780136893,
          "replace_card": 0.00011768836510703695
        },
        "choice": "freeze_card",
        "confidence": 0.9514991393911747
      },
      "unknown_charge": {
        "type": "noul",
        "noul": 0.141568209130773
      },
      "team": {
        "type": "choice",
        "probabilities": {
          "cards": 0.8397145103926745,
          "fraud": 0.13897325126695045,
          "payments": 0.021312238340375097
        },
        "choice": "cards",
        "confidence": 0.7595717655890116
      }
    }
  }
}
Fig. 3 — Each tab is one request: the input on the left, its typed questions on the right. Every bar is a probability Bongard-mini returned through the public demo on 30 September 2026, replayed here.

Judgments in motion.

In each recording the game engine describes the position and lists the legal moves. Bongard-mini returns a probability for every move, shown on the right, and the most likely move is played.

Plate 1 — 2048 · Game recording

The engine lists the legal moves with a short lookahead verdict for each; Bongard picks one. This game reaches the 4096 tile and scores 55,140. Nine of twelve games reach 2048.

03 — Results

Put the judgment to the test.

How often it is right, where larger systems still lead, and whether its probabilities and its speed hold up in use.

DecisionBench scores 23,900 decisions from 43 tasks. Bongard-mini places fourth of 61 systems, ahead of Jev 1.13, DeepSeek V4.1 Flash and GPT-5.6 Luna.

Judgment across tasks

23,900 decisions · 43 tasks · eval split

Selected systemsAccuracy
Imajev-4B79.7%
Bongard-mini78.05%
Winnow-12B76.7%
Jev 1.1372.0%
DeepSeek V4.1 Flash71.0%
GPT-5.6 Luna69.9%
Table 1 — Bongard: 78.05% with the official runner, 30 September 2026; all 23,900 rows answered, no errors or truncation. On 20,959 rows without detected source-text overlap with training data: 77.73%. Serving limits: 262,144 tokens per request, 131,072 per sequence, up to 255 candidates. Run and protocol. Reference scores: DecisionBench, 29 September 2026.

04 — How it works

One reading, many decisions.

A T5 encoder reads the situation in both directions, with the question instructions in view. Separate decoder branches use that shared evidence. A trained head scores each question’s candidates and returns probabilities. No text is generated.

Noul
Is this statement true? The answer is the probability that it is.
Choice
Which named option applies? Up to 255 options, each with a description, and a probability for every one.
Score
Where on an ordered scale? Two to ten levels, a distribution over them, and its expected level.

T5Gemma 2 · 4B-4B encoder–decoder · 7.51 billion parameters · one encoder pass, questions processed in groups of up to eight

A wrapper changes the interface. Bongard trains the judgment.

Jev, Kev and Bongard share a class of questions. How a model learns and reads the evidence determines the route it takes to a judgment.

SystemEvidenceLearningReadout
BongardT5 encoder–decoder. Shared, bidirectional evidence; separate question branches.Full text-model training on judgments, semantic relationships and action outcomes.A trained head scores the supplied candidates.
JevClosed architecture. Its API accepts shared evidence and multiple questions.RLCD, as described by TypeSafe. The full recipe is private.Direct probabilities through the System One API.
KevQwen causal backbone. The server reuses a state cache across questions.LoRA adapters and a pointer head trained on labeled decisions.A trained pointer head scores each option against a decision state.
OpenJev defaultPretrained DiffusionGemma. State and questions share one input.An inference wrapper around existing weights.Answer-token probabilities at masked positions.

Published designs, September 2026. RLCD: Reinforcement Learning for Calibrated Decisions. See the report for sources and additional implementations.

Read the whole situation.

A clause changes the meaning of an earlier clause. A log entry explains an earlier failure. Bidirectional encoding lets those facts shape each other’s representation before the judgments are read.

Put experience into the model.

Semantic pairs teach relationships across expressions and inputs. Sandbox outcomes teach what follows an action. These learning goals shape the evidence and the judgment together.

Choose where to spend capacity.

T5 separates reading a state from judging each question. Bongard uses balanced 4B-4B stacks today. This structure also supports future designs with more capacity for shared reading and less for each question.

Full-model training takes more data and compute than adapter training. Bongard makes that investment to develop a reusable judgment model. Its natural workload is a situation with related evidence and several decisions: a document, an incident or an agent’s current state.

05 — Training

Intuition is learned.

Experience shapes a judgment before the moment it is needed. Bongard learns in three stages: how to judge, how meanings relate, and what follows an action.

  1. 01

    Judgment

    Learn to judge from text, records and images. Paired examples teach which changes should alter an answer and which should leave it unchanged.

  2. 02

    Relationships

    Learn semantic relationships across words, images and states. A joint-embedding objective connects a judgment to the content and outcomes it concerns.

  3. 03

    Outcomes

    Learn from what actions produce. Sandbox rollouts and exact solvers supply outcome distributions for candidate actions. The model trains directly on those distributions.

What training adds to T5Gemma 2

Accuracy · pretrained models as published

BenchmarkGemma 3 4BT5Gemma 2Bongard
BoolQpublished 0-shot75.5%79.3%89.2%
WinoGrandepublished 5-shot69.9%71.6%85.7%
SocialIQApublished 0-shot49.8%49.9%60.9%
Table 4 — Bongard-mini answers zero-shot in one forward pass. T5Gemma 2 4B-4B and Gemma 3 4B: T5Gemma 2 report, Table 4.

What training changes

EvaluationBeforeAfter
Accuracy on held-out rephrasingsFull stage 275.7%85.9%
Accuracy on the frozen sandbox panelStage 3 · 16,992 questions at 1,953 states50.6%64.8%
Model-routing accuracyStage 2 → final · 600 held-out RouterBench items56.2%62.5%
Table 5 — Each row compares checkpoints on the named evaluation. The report gives the training conditions, controls and full results.

06 — Release

Open weights, open report.

Bongard-mini is the final model from the three-stage training program. Its weights are released under the Gemma Terms of Use. The Apache 2.0 code covers training, calibration and serving.

Model
Bongard-mini — 7.5 B parameters, encoder–decoder 4B + 4B, built on T5Gemma 2
Try it
Hugging Face Space
Weights
AgentBull/bongard-mini on Hugging Face
Report
Bongard: Training Machine Intuition (PDF)
Code
AgentBull/bongard on GitHub
License
Gemma Terms of Use
Authors
Li Ding, Haidi Jin, Chen Ji
Published
AgentBull Pte Ltd, Singapore, September 2026

# hf download AgentBull/bongard-mini --local-dir bongard-mini
from bongard.inference import Predictor

model = Predictor.load("bongard-mini")

model.predict({
    "state": {"subject": "Charged twice for order #4821"},
    "questions": {
        "refund": {"type": "noul",
                   "instructions": "Is the customer asking for a refund?"},
        "route": {"type": "choice",
                  "instructions": "Which team should handle this?",
                  "criteria": {"billing": "", "shipping": "", "technical": ""}},
    },
})

Cite

@techreport{ding2026bongard,
  title  = {Bongard: Training Machine Intuition},
  author = {Ding, Li and Jin, Haidi and Ji, Chen},
  institution = {AgentBull Pte Ltd},
  year   = {2026},
  month  = sep
}