Life sciences · Preprint
arXiv · September 8, 2026
Early or partial results. Treat as a signal, not a conclusion.
CoVeR is a new deterministic token-pruning algorithm for multi-view 3D reasoning in vision-language models that uses spatial coverage to select visual tokens without learned signals. In computational benchmarking on three 3D reasoning tasks, it reportedly retains 93.5% of full-token performance using only 8% of visual tokens, outperforming prior methods by 3.9 percentage points on average. This is a methods paper without peer review, clinical validation, or evidence of impact on real-world applications.
Preprint. Intervention: CoVeR: a deterministic, training-free token selector using spatial coverage that operates on token coordinates without learned signals. Compared with: Learned importance methods (attention/encoder-based token ranking) and voxelization methods; prior state-of-the-art approaches on 3D reasoning benchmarks.
CoVeR preserves 93.5% of full-token performance with approximately 8% of visual tokens CoVeR outperforms prior state-of-the-art methods by 3.9 percentage points on average across three 3D reasoning benchmarks Generalises as a plug-and-play module across four vision-language models
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
A novel computational method for token pruning in vision-language models, demonstrated on 3D reasoning benchmarks with no peer review and no clinical or translational validation reported.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sidestepping the scarcity of annotated 3D data. However, it produces thousands of redundant visual tokens whose cost grows with every view. Existing visual token pruners fall into two families, each limited in the 3D multi-view setting. Learned importance methods rank tokens by attention or encoder features; because redundancy here is fundamentally spatial, they keep near-duplicate tokens from a few prominent regions and leave most of the scene unrepresented. Voxelization methods improve spatial coverage but cannot enforce an exact token budget and saturate as multi-view observations overlap in 3D, capping retention well below the target. We show that spatial coverage is associated with 3D reasoning performance and introduce CoVeR, a deterministic, training-free selector that uses only token coordinates, with no learned signals. CoVeR selects tokens that collectively cover every region of the scene, and solves the limitations of both families: it enforces an exact per-scene budget, breaks the voxelization saturation plateau, and avoids the near-duplicate selections of learned importance. Extensive experiments show CoVeR outperforms prior SOTAs on all three 3D reasoning benchmarks and generalizes as a plug-and-play module tested across four VLMs. Notably, with only $\approx$8% of visual tokens, it preserves 93.5% of full-token performance, surpassing SOTA by 3.9 percentage points on average across benchmarks.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.