Bahlul SI · The Bangla test bench
Tested in Bangla, on your documents, before it goes live.
General benchmarks say little about your circulars, your scanned forms or your customers' mixed Bangla and English. So every Super Intelligence (SI) system Bahlul World builds is scored on questions written from your own documents, by two native Bangla reviewers, against a pass mark you approve.
“Will it give right answers on our documents, in Bangla?”Head of operations
Your situation
Why a public benchmark is not your answer
Language models are scored on public tests, mostly in English. Those scores say nothing about a scanned 2014 circular, a fee table in Bangla digits, a question typed half in Bangla and half in English, or a request the system should refuse. The Bangla test bench is Bahlul World's method for scoring any SI system on the documents it will serve. It is part of every SI proof, every server acceptance and every pilot, and the full set runs every month under managed SI operations.
Build
About 200 questions from your own documents
In the first week of an SI proof we write about 200 test questions with your subject experts, from your real documents, help-desk logs and emails. Real wording stays: spelling mistakes, Bangla typed in Latin letters, the way people actually ask. Personal data is masked. Each question records its expected answer and the source document with its page.
At least 50 are marked critical: a wrong answer there could cost money, break a rule or mislead a customer or citizen. About 50 are held out and never seen by the engineers who tune the system; they are scored only in the final round. The test set is yours. It stays in Bangladesh, and you keep it to test any system, from us or from anyone else.
The six categories
What the questions cover
| Category | Of 200 questions | What a pass looks like |
|---|---|---|
| Answer in the documents | 70 | The right, complete answer, with the right source shown |
| No answer in the documents | 30 | Says plainly that the documents do not cover it, does not guess, points to the right office |
| Mixed Bangla and English | 25 | Answers the question asked, in the user's language, including Bangla typed in Latin letters |
| Scanned pages with broken text | 25 | The right answer, or a clear note that the page cannot be read |
| Numbers and dates | 30 | Every figure and date exactly right: Bangla and Latin digits, lakh and crore, both calendars |
| Requests it must refuse | 20 | Refuses politely, uses nothing the person may not open, points to the proper channel |
Proposed counts for a 200-question set, with at least 20 in each category that applies. If your documents hold no scans, those questions move to other categories and the report says so.
The scorecard
Six categories, one pass mark
The six categories, their proposed share of a 200-question set, and the proposed pass mark of 85% on the critical questions.
Score
Two native Bangla reviewers, each on their own
Two reviewers score every answer: one is your subject expert, the other a native Bangla reader outside the team that built the system. Neither sees the other's scores or knows which version produced the answer. Each answer gets 2 (right, with the right source), 1 (partly right, no harm done) or 0 (wrong, invented, a guess where no answer exists, or a request carried out that should have been refused).
The two reviewers calibrate on 20 practice questions first. Afterwards their agreement rate is checked against a proposed minimum of 80%, differences are discussed and an agreed score recorded; any they cannot settle goes to your named approver, never to the build team. A critical failure is an agreed 0 on a critical question.
Decide
A pass mark you approve, before any scoring
Bahlul World proposes 85% or better on the critical questions and no critical failures. Your named approver confirms or changes it in writing on day 5 of the proof, before any answer is scored. Once scores are seen, the pass mark can only go up. For answers given to the public on fees, deadlines or entitlements, we recommend a higher bar, for example 95% on the critical set.
Meeting the pass mark is evidence for a decision, and the decision is yours: go, go with limits, or no-go. A weak category is a reason to go live with limits even when the pass mark is met. The same test, at the same or a higher pass mark, is the acceptance test for your server, and the balance for the set-up is due only when it passes.
Re-test
Run again after every change
The test runs again before any of these goes live: a new model or model version, a new Bahlul SI release, a change to prompts or refusal rules, a change to search or OCR settings, a move to new hardware, or a change to more than about a tenth of the document set. A small change needs the critical set at least; a model change needs the full set. Under managed SI operations the full set runs every month, and the report lists every question that passed before and fails now.
The report
What a report states, and what it never states
Each round ends with a one-page report for your approver: the system version and the date it was frozen; the pass rate by category and on the critical set; every critical failure, what went wrong and the fix planned; the reviewers' agreement rate; and every question that passed before and fails now. The go-live round's results go into the system card you keep.
A report states how one version scored on one test set on one date. It never calls a system accurate, safe or compliant in general, and it never implies a certification. We quote no accuracy figure for Bahlul SI other than your own result against your own pass mark.
The first step
Where the test bench starts
The test bench is part of the three-week SI proof on your documents, BDT 5.5 lakh, a proposed 2027 price before VAT. It ends with your scores, the server sized, a business case and one recommendation: buy the server, run a larger pilot, or stop. If the result misses your pass mark and the gap cannot be closed soon, the recommendation is to stop, and we say so.
FAQ
Questions we are asked
Who writes the questions?
Our delivery lead with your subject experts, from your real documents, help-desk logs and emails, in a two-hour workshop in week one. Your named approver signs off the set, the critical questions and the pass mark before any scoring.
Who are the reviewers?
One is your own subject expert. The other is a native Bangla reader outside the build team, from your organisation or ours. Both sign a confidentiality undertaking.
What happens if it fails?
A failed round is fixed within Bahlul SI's settings and run again. If the final round still misses your pass mark, the recommendation is to stop or to fix the documents first, and you keep the test set and the report.
Does the test cover voice?
In version 1 a voice system is tested on its transcribed text. Speech and handwriting categories join the test bench in February 2028, and voice goes live only where it passes.
Write the first twenty questions with us
Bring ten documents your team searches most. In the SI briefing we show Bahlul SI answering from them; in the proof, your questions decide.