Serious about
a silly test.
A pretty bird isn’t enough. We score the relationship between a pelican, a bicycle, and the act of riding.
20 points for the SVG. 80 for the picture.
The Machine Score evaluates the file’s technical health. The Vision Score evaluates the standardized 1024 × 1024 PNG. Their sum is rounded to an integer from 0 to 100. Decorative polish never replaces semantic correctness.
Machine Score / 20
| Check | Points |
|---|---|
| Valid SVG parsing | 2 |
| Successful sanitization | 2 |
| Successful rendering | 3 |
| Non-blank, meaningful pixels | 3 |
| Content inside the viewport | 2 |
| Canvas utilization | 2 |
| Reasonable structural complexity | 2 |
| No prohibited or unsupported content | 2 |
| Technical SVG quality | 2 |
These checks are deterministic for the same input and renderer. Visible-pixel bounds and edge density are heuristics: invisible geometry outside the viewport cannot be reliably inferred. Unsupported elements or attributes are removed and cost points in the prohibited-content check.
Vision Score / 80
| Dimension | Points | What the judge looks for |
|---|---|---|
| Pelican | 15 | Recognizable bill, pouch, and coherent anatomy. |
| Bicycle | 15 | Two wheels, connected frame, seat, bars, and pedals. |
| Riding relationship | 25 | Body on the seat, feet related to pedals, sensible contact. |
| Mechanical plausibility | 15 | Wheel alignment, connected frame, fork and drivetrain. |
| Composition | 10 | Readable subject, clear spatial relationships, framing. |
One image. More than one opinion.
Development defaults to one judge run; production defaults to two. If two totals differ by 5 points or fewer, we average each dimension. Their explanations come from the lower-scoring run. If the totals differ by more than 5, a third run breaks the tie: we use the complete result with the median total. All raw responses are retained. Humorous comments never change scores.
Safety is part of the pipeline.
Raw uploads are private server-side assets. We reject document types, custom entities, malformed XML, excessive size and complexity. We remove scripts, event handlers, embedded HTML, images, external resources and unsupported SVG content. Safe static SVG is rendered in a separate process with a deadline, a JavaScript heap budget, fixed white background and a bundled font. Public pages serve PNG only.
Versions stay attached to results.
Each submission retains scoring version pelican-v1, machine-v1, prompt pelican-judge-v1, judge provider and model, renderer version, input and render hashes, and timestamps. Changing an engine version creates a new cache identity and result instead of overwriting previous scores. The public leaderboard compares only pelican-v1.
The judge has limitations.
Pelican Score is an experimental community benchmark. AI judges can miss details, favor familiar visual styles, misread contact points, or disagree. Confidence is the judge’s own estimate, not a calibrated probability. A single prompt does not measure a model’s general intelligence. Model names are self-reported; scores are not independently verified model benchmarks.
Without a configured API key, development uses a fixed placeholder labeled Development Mock Score. It performs no visual recognition, is unlisted, and never enters the official leaderboard. Production fails closed until a real judge is configured.
Try the test ↗