Why semi-visible jets resist a kinematics-only search
Semi-visible jets arise in Hidden Valley models where a confining dark sector showers and hadronizes, producing a mixture of Standard Model hadrons and stable dark hadrons inside the same jet cone; the invisible component escapes as missing transverse momentum correlated with visible jet activity [1][5]. Because the dark sector introduces poorly constrained parameters such as the dark hadron mass, the invisible fraction r_inv, and the dark hadronization scale, a single benchmark point cannot capture the full signature range, which is why the anchor paper scans 18 signal benchmarks across two mediator masses [1]. Earlier work established the basic phenomenology: jet substructure observables differ between semi-visible and light-quark/gluon jets, but the comparison is sensitive to the assumed dark shower model and to the dark hadron fraction [7]. The Snowmass 2021 dark showers report further warned that infrared-and-collinear-unsafe substructure variables used in Hidden Valley searches must be validated in control regions, and that the unknown dark hadronization scale can strongly reshape substructure observables [5].
This is the gap the new paper targets. Existing ATLAS and CMS SVJ searches have largely exploited global event and leading-jet observables, with comparatively little attention to jet substructure or MET-correlated angular variables [1]. The anchor paper's explicit motivation is that the visible radiation pattern of a semi-visible jet may retain discriminating information even when the missing-energy signature weakens, for example at small invisible fraction [1].
Five classifiers, one dataset: what the controlled comparison isolates
The anchor paper trains five deep-learning classifiers on a common balanced dataset of 300,000 SVJ signal events and 300,000 QCD background events, using identical event selection, training/validation splits, and optimization protocol so that performance differences can be attributed to representation and architecture rather than training configuration [1]. The three standalone representations are a Vision Transformer on 50x50 Lund Jet Plane images, a JetLOV network on the C/A reclustered tree, and a 15-dimensional MLP vector of global and jet-level observables; two fusion networks combine the image or tree branch with the MLP branch [1]. Events are generated with PYTHIA 8.3.17, passed through DELPHES 3.5.1 with the CMS detector card, and jets are reconstructed with anti-kT R=0.8 [1]. Three additional benchmarks at an intermediate mediator mass of 3.5 TeV, not used in training, test interpolation across the dark sector parameter space [1].
The headline ordering is unambiguous. The MLP reaches accuracy 0.9054 and AUC 0.9653, JetLOV reaches 0.9003 and 0.9572, and the ViT reaches 0.8629 and 0.9293 [1]. The two fusion networks dominate: ViT+MLP reaches 0.9389 and 0.9846, and JetLOV+MLP reaches 0.9428 and 0.9872 [1]. Because the ViT and JetLOV consume the same declustering sequence, the gap between them is attributable to representation rather than physical input: the tree preserves the ordering and branching relations of splittings, while the LJP image bins each splitting independently and is permutation-invariant across splittings [1].
Why the hierarchical tree beats the Lund-plane image
The anchor paper's interpretation is that a non-negligible fraction of the discriminating information in a semi-visible jet is encoded in the correlated sequence of splittings rather than in the distribution of individual splittings alone [1]. This is physically motivated: the invisible component removes energy at intermediate stages of the dark shower, so the visible remnant is characterized not only by modified splitting scales but by their correlated branching history [1]. The image representation is also statistically sparse, with 2,500 pixels for typically O(10) resolved splittings, forcing the image encoder to process many empty pixels; JetLOV achieves its performance with roughly one third of the image encoder's parameters [1].
This finding sits alongside a broader jet-tagging literature in which graph and tree representations have been argued to offer improved interpretability and reduced simulation dependence. The CMS thesis on Hidden Valley searches reports that a LundNet graph neural network delivers excellent BSM jet-tagging performance with improved interpretability and reduced Monte Carlo dependence, and that its application to SVJs is the first of its kind [2]. A separate explainability study of LundNet, ParticleNet, and Particle Transformer finds that all three architectures rediscover canonical QCD substructure observables, with the Particle Transformer approaching functional equivalence with tau21 and tau32 in regimes where those observables are well resolved [3]. That study also notes that moderate correlation values leave room for non-linear effects or genuinely novel features beyond classical observables [3].
The fusion gain is largest where searches need it most
The anchor paper reports that the relative advantage of the fusion networks becomes increasingly pronounced as background rejection is tightened. At 80% background rejection, JetLOV+MLP improves signal efficiency by about three percentage points relative to the MLP; at 99% rejection, the improvement reaches fifteen percentage points, with signal efficiency 0.819 for JetLOV+MLP versus 0.643 for the MLP and 0.508 for the ViT [1]. The paper's explanation is that remaining QCD background events at high rejection can fluctuate toward the large multiplicity and diffuse radiation patterns characteristic of semi-visible jets, making them resemble signal in low-dimensional observables, while their detailed splitting structure remains governed by ordinary soft-collinear QCD radiation that the hierarchical tree can resolve [1].
This is the regime that matters for real searches. The ATLAS semi-visible jet search uses a Particle Flow Network on track-level inputs and an anomaly-detection tool, and excludes Z' masses between 2000 and 3200 GeV at 95% CL for r_inv from 0.2 to 0.37 with 140 fb^-1 [6]. The CMS thesis reports first collider searches for lepton-enriched SVJ final states using LundNet and a DNN-based background estimation, excluding Z' mediator masses up to about 4.7 TeV for SVJl and about 4 TeV for SVJtau [2]. The anchor paper's fusion result suggests that the substructure information these searches already partially exploit could be combined more systematically with global kinematics.
What the comparison does not yet settle
The anchor paper's conclusions rest on simulated Z'-mediated Hidden Valley benchmarks with PYTHIA 8.3.17 and DELPHES 3.5.1, not on real LHC data, and the authors explicitly note that further studies are needed to assess dependence on simulation modeling and training fluctuations [1]. The Snowmass report documents that hadronization parameters in the PYTHIA Hidden Valley module remain uncertain and that changes in these parameters can shift jet substructure observables, which limits the interpretability of substructure-based results [5]. The anchor paper also acknowledges that its study does not enable direct comparison between different benchmark points [1].
A competing approach in the literature is unsupervised anomaly detection. A detailed VAE study finds that mass-decorrelated VAEs almost lose the ability to act as anomaly detectors, and that a semi-supervised outlier-exposure variant is needed to recover sensitivity while maintaining mass decorrelation [4]. That work shows that purely unsupervised learning without guidelines may not give the optimal solution for jet tagging [4]. The anchor paper's supervised, benchmark-specific classifiers and the anomaly-detection approach answer different questions: the former quantifies achievable discrimination for known SVJ topologies, while the latter aims for model-independent sensitivity. How the two would perform on the same events, and how the fusion gain would survive detector systematics and background estimation uncertainties, remains open.
About These Sources
This research page is built on 7 studies (3 peer-reviewed, 4 preprints) — published from 2021 to 2026, 4 from 2024 or later, 3 in Q1 journals, collectively cited 133 times — selected as the most relevant from 13 studies that passed quality screening, drawn from 148 papers retrieved from a database of over 500 million.
Sources used in this answer
Hunting the Unseen: Deep Learning Analysis for Semi-Visible Jet Tagging
The anchor paper trains five deep-learning classifiers on Z'-mediated Hidden Valley benchmarks and finds that global kinematic observables outperform both Lund Jet Plane images and hierarchical tree representations, while fusing the tree with global observables yields the best overall discrimination with AUC 0.9872.
Listening to the echoes of a Hidden Valley at the LHC with the CMS experiment
The CMS thesis presents first collider searches for lepton-enriched semi-visible jet final states using LundNet and a DNN-based background estimation, excluding Z' mediator masses up to about 4.7 TeV for SVJl and about 4 TeV for SVJtau.
Explainable AI for Jet Tagging: A Comparative Study of GNNExplainer, GNNShap, and GradCAM for Jet Tagging in the Lund Jet Plane
An explainability study of LundNet, ParticleNet, and Particle Transformer finds that all three architectures rediscover canonical QCD substructure observables, with the Particle Transformer approaching functional equivalence with tau21 and tau32 in well-resolved regimes.
Variational autoencoders for anomalous jet tagging
A VAE study for anomalous jet tagging finds that mass-decorrelated VAEs almost lose anomaly-detection capability, and that a semi-supervised outlier-exposure variant is needed to recover sensitivity while maintaining mass decorrelation.
Theory, phenomenology, and experimental avenues for dark showers: a\n Snowmass 2021 report
The Snowmass 2021 dark showers report documents that PYTHIA Hidden Valley hadronization parameters remain uncertain, that changes in these parameters can shift jet substructure observables, and that IRC-unsafe variables in Hidden Valley searches must be validated in control regions.
Search for new physics in final states with semivisible jets or anomalous signatures using the ATLAS detector
The ATLAS semi-visible jet search uses a Particle Flow Network and an anomaly-detection tool on 140 fb^-1 of 13 TeV data, excluding Z' masses between 2000 and 3200 GeV at 95% CL for r_inv from 0.2 to 0.37.
Exploring jet substructure in semi-visible jets
A jet substructure study compares semi-visible jets and light-quark/gluon jets using the PYTHIA Hidden Valley module across different dark hadron fractions, finding that substructure observables differ but that the comparison is sensitive to the assumed dark shower model.
