Life sciences · Preprint
arXiv · September 8, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint introduces Per-Variable Surgery (PV-Surgery), an optimizer-side training strategy that addresses gradient disagreement across variables in multivariate time-series forecasting. The authors report average reductions in MSE of 3.61% and MAE of 2.93% across five backbones and seven datasets, based on observations that 30.6% of pairwise variable-wise gradient cosine similarities are negative. However, the work is unreviewed and lacks comparison to established forecasting baselines or statistical significance testing.
Empirical computational benchmarking study. Multivariate time-series forecasting tasks across seven unnamed datasets evaluated on five deep learning backbones (specific architectures not named in abstract).. Intervention: Per-Variable Surgery (PV-Surgery): an optimizer-side training strategy using variable-wise gradient proxies, reliability-aware layer selection, conditional pooling, and common-direction surgical gradient alignment.. Compared with: Baseline shared training (mean-loss optimization); specific competing methods not named in abstract..
30.6% of pairwise cosine similarities between variable-wise gradients are negative on average across seven datasets Under shared training, 35 of 64 variables perform worse than a full-input single-target oracle PV-Surgery lowers MSE by 3.61% and MAE by 2.93% on average across five backbones, seven datasets, and four horizons
Harmed variable fraction is not reliably predicted by gradient conflict frequency
Not applicable. This is a computational optimization technique for forecasting algorithms. Clinical or operational relevance to practice depends on whether forecasting performance improvements translate to meaningful outcomes in downstream applications.
This is a preprint describing an algorithmic optimization technique for multivariate time-series forecasting, evaluated empirically across multiple datasets and backbones but without peer review, clinical validation, or comparison to established baselines.
As stated by the source record.
Quoted from the source exactly as published.
Not applicable. This is a computational optimization technique for forecasting algorithms. Clinical or operational relevance to practice depends on whether forecasting performance improvements translate to meaningful outcomes in downstream applications.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
In data-driven training, multivariate time-series forecasting is usually optimized with a scalar loss averaged over samples, variables, and horizons. This averaging is convenient, but the optimizer sees only the aggregated gradient, which does not reveal whether the variable-wise contributions align or oppose one another. To quantify how often this disagreement arises, we measure the variable-wise gradients directly and find that 30.6% of their pairwise cosine similarities are negative on average across seven datasets. However, conflict and harm are not the same thing. Under shared training 35 of the 64 variables do worse than a full-input single-target oracle, and the harmed fraction is not reliably predicted by how often gradients conflict. We propose Per-Variable Surgery (PV-Surgery), an optimizer-side training strategy for backbones with cache-compatible layers. One backward pass builds variable-wise gradient proxies from output-side signals and keeps the pointwise forecasting loss. Reliability-aware selection targets layers whose proxy sums closely approximate their shared-gradient slices. Conditional pooling forms anchor and conflict pools without dropping variables. Common-direction surgery aligns variable or pooled gradients with their normalized mean and restores input norms to avoid reweighting. In experiments across five backbones, seven datasets, and four horizons, PV-Surgery lowers MSE by 3.61% and MAE by 2.93% on average. For multivariate forecasting, this indicates that the variable-wise structure hidden by mean-loss training is a usable optimization signal.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.