Proof · how we test
Tested on signatures it has never seen.
We test nib v2.2 on public signature datasets, straight out of the box: no training on the test data, one policy for everyone. Results will be published once the dataset owners have agreed.
Latin 55 writers · sealed set
Devanagari 160 writers · sealed set
Bengali 100 writers · development set
RESULTS TO FOLLOW
01 · What we measure
What happens to a skilled forgery.
A skilled forgery is made by someone who studied the real signature. It is the hard case, and the one that matters in fraud, so it is what we measure first.
Of the forgeries made by someone who studied the real signature, how many the default policy would approve.
The point where false accepts and false rejects are equally likely. Lower is better.
How many real signatures go straight through, and how many are sent to a person.
02 · More references
Tested with few references, and with many.
For each writer we draw 3, 5 and 10 reference signatures at random, then check the rest. 10 is the most a check uses. Each run is repeated with 3 different draws.
- references 3
- references 5
- references 10
03 · The honest part
We’d rather ask a person than wave a forgery through.
Nib is tuned to be strict. The default policy approves only from 0.95; anything less sure goes to a person. Nothing in FLAG or FAILED is approved without a person.
- APPROVEfiled
- FLAGto a person
- FAILEDto a person
04 · How we test
Rules we hold ourselves to.
No training on any benchmark signature. One default policy for every dataset, never tuned per set.
CEDAR and Hindi are sealed: we only ever see headline summaries, never individual results.
Bengali is our development set. We say so, so you can weigh its results accordingly.
Each sealed run is recorded with its date and version. CEDAR’s flags were viewed once, before these rules existed.
| Dataset | Script | Writers | Role |
|---|---|---|---|
| CEDAR | Latin | 55 | sealed |
| BHSig260-Hindi | Devanagari | 160 | sealed |
| BHSig260-Bengali | Bengali | 100 | development |
The fine print
nib v2.2 · default policy (default@1). Each run averages 3 random draws of reference signatures per writer. Definitions and limits are in the whitepaper.
Results will be published once the dataset owners have agreed.