Life sciences · Preprint
arXiv · August 18, 2026
Early or partial results. Treat as a signal, not a conclusion.
This unreviewed computational study examines whether surgical skill assessment models learn representations that generalize across different scoring rubrics (GOALS and OSATS). Results show asymmetric transfer: models pretrained on JIGSAWS transfer effectively to LASANA (CCC 0.77–0.80), but transfer to JIGSAWS fails across all approaches, suggesting annotation inconsistencies and that task-specific prediction heads dominate over transferable backbone features.
Computational transfer learning analysis with controlled experiments and ablations. Intervention: Backbones pretrained on JIGSAWS, ASAM, self-supervised pretraining, contrastive learning. Compared with: End-to-end supervised training; Kinetics-pretrained backbone controls.
Backbones pretrained on JIGSAWS achieve CCC values of 0.77 to 0.80 on LASANA, closely matching end-to-end baseline Transfer to JIGSAWS fails across all evaluated methods (end-to-end, ASAM, self-supervised, and contrastive learning) Task-specific heads carry the majority of skill prediction burden; backbone contribution limited to spatiotemporal features
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
An unreviewed exploratory analysis of cross-domain transfer in surgical skill assessment models, using controlled experiments to probe representation learning but without clinical validation or a definitive answer to its central question.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Vision-based surgical skill assessment has shown strong in-domain results, yet a fundamental question remains unasked: do these models learn transferable representations of surgical proficiency, or do they merely encode dataset-specific visual patterns? This paper systematically analyzes what limits cross-domain skill transfer between the GOALS and OSATS assessment scales using the LASANA and JIGSAWS datasets. Each evaluated method serves a targeted diagnostic purpose: end-to-end training to test whether supervised skill learning transfers directly, Adaptive Sharpness-Aware Minimization (ASAM) to probe whether flatter loss landscapes improve generalization, and augmentation-based self-supervised and contrastive learning to assess whether domain-invariant pretraining decouples skill from visual context. Transfer is evaluated in both directions using a disjoint-participant held-out test set for JIGSAWS. Results reveal an asymmetry: backbones pretrained on JIGSAWS achieve CCC values of 0.77 to 0.80 on LASANA, closely matching the end-to-end baseline, showing cross-rubric transfer is feasible when the target domain provides consistent supervision. Transfer to JIGSAWS fails across all methods, likely due to annotation inconsistencies. Control experiments with a Kinetics-pretrained backbone suggest task-specific heads carry the majority of the skill prediction burden, while the backbone need only provide adequate spatiotemporal features. These findings offer a new perspective on vision-based skill assessment: the central question of whether skill representations transfer across scoring systems has not been previously investigated. Results indicate the visual component is dominant but not solely responsible for skill prediction; further work is needed to conclusively disentangle transferable skill features from those bound to a specific visual domain.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.