A result only matters if it holds up. We benchmark under the hardest realistic conditions: scaffold-aware splits so test molecules are structurally distinct from training, multi-seed protocols to separate real signal from lucky runs, and honest baselines we genuinely try to beat.
We report effect sizes and confidence intervals rather than single cherry-picked numbers, and we release code so others can reproduce what we claim. If a method only wins under a favourable split, it hasn’t won.