Life sciences · Preprint
arXiv · September 4, 2026
Posted before peer review. The findings may change or fail to hold.
WEECFP-SuRGE is a novel deep learning architecture for molecular property prediction that demonstrates competitive performance on standardized computational benchmarks without external pretraining. The method is presented as a technical advance in molecular fingerprinting and transformer design but has not undergone peer review and contains no clinical validation, biological experiments, or comparison to experimental ground truth.
Preprint. Molecular datasets (TDC ADMET suite, MoleculeNet); no human subjects. Intervention: WEECFP-SuRGE transformer architecture combining WEECFP molecular fingerprints with SuRGE (Substructure Rotary Graph-distance Encoding) applied to self-attention. Compared with: Classical fingerprint baselines, MapLight+GNN, and other methods on TDC ADMET and MoleculeNet benchmarks.
WEECFP-SuRGE Blend ranked #2 overall on TDC ADMET leaderboard and #1 among non-pretrained methods Achieved leaderboard #1 finishes on 5 of 22 benchmarks: Pgp, Lipophilicity, CYP2D6 Substrate, Clearance Microsome, and LD50 WEECFP tokenization reconstructed exact canonical SMILES in 99.9% of in-distribution molecules across 9 MoleculeNet datasets
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is an unrefereed technical preprint introducing a machine learning method for molecular fingerprinting with competitive leaderboard performance but no peer review and no clinical or translational validation.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
We introduce WEECFP, a parameter-free 1024-dimensional continuous molecular fingerprint that scatters each Morgan substructure across roughly thirty-two signed positions of a single vector, and WEECFP-SuRGE, a transformer architecture whose self-attention applies SuRGE (Substructure Rotary Graph-distance Encoding) -- a RoPE-like rotation parameterized by molecular shortest-path graph distance -- to WEECFP substructure tokens. A 7-model blend of this architecture (the WEECFP-SuRGE Blend) achieves the lowest average regression rank on the TDC ADMET leaderboard; is #2 overall on the TDC ADMET leaderboard (behind only pretrained MapLight+GNN), and is #1 overall among methods that use no external pretraining; takes leaderboard #1 finishes on Pgp, Lipophilicity, CYP2D6 Substrate, Clearance Microsome, and LD50 (with the WEECFP-NoSuRGE Blend separately reaching #1 on HIA) across the full 22-benchmark suite -- without any external pretraining. On MoleculeNet, WEECFP-SuRGE beats every classical-fingerprint baseline on 3 of 4 regression tasks (ESOL, Lipophilicity, QM9). We further show that WEECFP tokenization is near-lossless: a greedy overlap reconstruction recovers the exact canonical SMILES of 99.9% of in-distribution molecules across 9 MoleculeNet datasets and 98.93% of molecules in a cross-dataset holdout (HIV->Lipophilicity), and that a three-reference farthest-first encoding of graph distance correlates at Pearson r = 0.901 with the true pairwise distance, enabling O(S) positional memory at matching accuracy.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.