MathGen

Revealing the Illusion of Mathematical Competence through Text-to-Image Generation

MathGen covers seven mathematical domains with representative prompts and reference illustrations.

MathGen spans counting, angles, fractions, functions, plane geometry, sets, and solid geometry. Each problem asks a text-to-image model to render a visually valid mathematical structure, not merely a plausible-looking picture.

Introduction

Modern generative models can solve increasingly difficult mathematical problems in text, but many real use cases require answers to be expressed visually through diagrams, plots, geometric constructions, and symbolic layouts. MathGen asks a direct question: does mathematical competence persist when the answer must be rendered as an image?

We introduce MathGen, a benchmark of 420 curated problems across seven core mathematical domains. The benchmark contains 350 Clean-Scene problems and 70 paired Open-Scene problems. Each task is evaluated with a Script-as-a-Judge protocol: problem-specific executable checks verify numerical, geometric, structural, and logical constraints deterministically.

420
Curated Problems

350 Clean-Scene problems plus 70 paired Open-Scene problems.

7
Mathematical Domains

Counting, angles, fractions, functions, plane geometry, sets, and solid geometry.

95.4%
Human Agreement

Script-based judgments align strongly with graduate-student human annotations.

Benchmark Design

MathGen separates controlled diagrammatic math from realistic open-scene math. This paired design helps diagnose whether failures come from the core mathematical constraint, from scene complexity, or from both.

Prompt. A problem describes the exact mathematical structure the model must render.
Generation. Each model produces one image per prompt under its default inference setup.
Verifier. Scripts check contours, lines, OCR, object counts, topology, color regions, or geometric relations.
Decision. A sample is correct only when all problem-specific constraints are satisfied.
Overview of the MathGen benchmark and evaluation pipeline.

Script-as-a-Judge provides transparent, reproducible, fine-grained verification of mathematical correctness.

Counting

Exact counts and attribute-based counts for target objects.

Angles

Angle measures, angle relations, and geometric angle construction.

Fractions

Fraction grids, proportional mappings, and ratio-preserving visual regions.

Functions

Coordinate axes, continuous curves, piecewise plots, and function relations.

Plane Geometry

Intersections, composite figures, and 2D geometric constructions.

Sets

Set operations, relations, membership, overlaps, and disjointness.

Solid Geometry

3D shapes, coordinates, projections, visibility, and occlusion.

Open-Scene Math

Matched clean problems embedded in richer real-world visual contexts.

Clean-Scene Leaderboard

Main results on the 350-problem Clean-Scene set. Accuracy is reported for each domain with 50 problems per domain. Closed-source models lead, but even the strongest systems remain far from reliable mathematical rendering.

53.7%
Nano Banana Pro

Best overall result, with at least 40% accuracy in every domain.

39.1%
GPT-Image-1.5

Strong on counting, fractions, and plane geometry, but weaker on sets, functions, and angles.

11.1%
Best Open-Source Overall

FLUX-2 is strongest among evaluated open-source models, highlighting the current gap.

Model Counting Angle Fraction Function Plane Set Solid Overall
Diffusion Models
SD-3-Medium Diffusion0.00.00.00.06.00.06.01.7
SD-3.5-Medium Diffusion4.00.00.00.012.02.06.03.4
SD-3.5-Large Diffusion12.00.02.00.012.02.04.04.6
FLUX-2 Diffusion8.02.08.02.042.08.08.011.1
PixArt-Sigma Diffusion10.00.02.00.012.02.08.04.9
PixArt-XL-2 Diffusion0.00.00.00.014.00.04.02.6
HiDream-I1 Diffusion6.00.00.02.04.02.06.02.9
Qwen-Image Diffusion22.00.08.02.024.04.06.09.4
Z-Image-Turbo Diffusion8.00.08.00.016.02.014.06.9
Autoregressive Models
Infinity-8B AR6.00.04.00.018.02.08.05.4
GoT-R1-7B AR8.00.00.00.016.02.02.04.0
Unified Models
BAGEL Unified4.00.00.00.014.00.02.02.9
show-o2-1.5B Unified0.00.00.04.018.00.02.03.4
show-o2-7B Unified0.00.00.04.010.00.08.03.1
Janus-Pro-1B Unified0.00.00.00.014.00.02.02.3
Janus-Pro-7B Unified0.00.00.00.012.02.06.02.9
BLIP3o-4B Unified4.00.02.00.014.02.08.04.3
BLIP3o-8B Unified6.00.02.00.014.02.04.04.0
OmniGen2-7B Unified4.00.02.00.012.02.08.04.0
Closed-Source Models
FLUX-2-Pro Closed22.010.020.018.054.016.020.022.9
FLUX-Kontext-Pro Closed10.00.010.04.018.06.04.07.4
Seedream 3.0 Closed14.00.02.00.024.06.08.07.7
Seedream 4.0 Closed20.00.06.010.036.06.014.013.1
Ideogram v3 Turbo Closed10.02.00.00.020.02.08.06.0
Nano Banana Closed20.08.024.010.064.010.024.022.9
Nano Banana Pro Closed48.054.050.070.072.042.040.053.7
Imagen 4 Closed12.02.02.02.012.00.012.06.0
Imagen 4 Ultra Closed20.06.016.04.042.08.016.016.0
GPT-Image-1 Closed32.012.044.016.068.020.024.030.9
GPT-Image-1.5 Closed56.024.054.022.070.020.028.039.1

Findings and Analysis

Closed models lead, but remain limited.

Nano Banana Pro reaches 53.7% overall, while most open-source systems remain below 10%.

Progress is uneven across domains.

Models can be strong on counting or plane geometry while failing functions, angles, or set relations.

Open scenes add difficulty.

Realistic visual context often reduces exact mathematical correctness, even when the underlying constraint is unchanged.

Clean-Scene and Open-Scene comparison examples.

Matched Clean-Scene and Open-Scene examples show how realistic context stresses mathematical fidelity.

Qualitative comparison of text-to-image models on MathGen tasks.

Representative cases illustrate that visually convincing outputs may still violate exact constraints.

Error ratio of qualitative and quantitative failures across MathGen topics.

Error ratios distinguish qualitative structural failures from quantitative numerical or proportional failures.

Representative MathGen error examples.

Common failures include incorrect counts, broken proportions, imprecise angles, invalid plots, and wrong set structure.

BibTeX

@misc{liu2026mathgen,
  title={MathGen: Revealing the Illusion of Mathematical Competence through Text-to-Image Generation},
  author={Liu, Ruiyao and Shen, Hui and Zhang, Ping and Hsieh, Yunta and Zhang, Yifan and Xu, Jing and Han, Qi and Li, Junchen and Lu, Jiawei and Ma, Jianing and Mo, Jiaqi and Chen, Sicheng and Zhang, Zhen and Wan, Zhongwei and Xiong, Jing and Wang, Xin and Liu, Ziyuan and Cao, Hangrui and Wong, Ngai},
  year={2026},
  eprint={2603.27959},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2603.27959}
}