<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">NHESS</journal-id><journal-title-group>
    <journal-title>Natural Hazards and Earth System Sciences</journal-title>
    <abbrev-journal-title abbrev-type="publisher">NHESS</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Nat. Hazards Earth Syst. Sci.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1684-9981</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/nhess-26-4529-2026</article-id><title-group><article-title>Infra-Net: a robust parallel decision-making network for discriminating natural hazards and anthropogenic infrasound events via multi-view feature learning</article-title><alt-title>Infra-Net</alt-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Li</surname><given-names>Hongru</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Li</surname><given-names>Xihai</given-names></name>
          <email>xihai_li@163.com</email>
        <ext-link>https://orcid.org/0000-0001-8348-8256</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Liu</surname><given-names>Jihao</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Luo</surname><given-names>Shengjie</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Zhang</surname><given-names>Yun</given-names></name>
          
        </contrib>
        <aff id="aff1"><label>1</label><institution>Rocket Force University of Engineering, Xi'an, 710025, Shaanxi, China</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Xihai Li (xihai_li@163.com)</corresp></author-notes><pub-date><day>23</day><month>September</month><year>2026</year></pub-date>
      
      <volume>26</volume>
      <issue>9</issue>
      <fpage>4529</fpage><lpage>4548</lpage>
      <history>
        <date date-type="received"><day>19</day><month>March</month><year>2026</year></date>
           <date date-type="rev-request"><day>13</day><month>April</month><year>2026</year></date>
           <date date-type="accepted"><day>11</day><month>September</month><year>2026</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2026 Hongru Li et al.</copyright-statement>
        <copyright-year>2026</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://nhess.copernicus.org/articles/26/4529/2026/nhess-26-4529-2026.html">This article is available from https://nhess.copernicus.org/articles/26/4529/2026/nhess-26-4529-2026.html</self-uri><self-uri xlink:href="https://nhess.copernicus.org/articles/26/4529/2026/nhess-26-4529-2026.pdf">The full text article is available as a PDF file from https://nhess.copernicus.org/articles/26/4529/2026/nhess-26-4529-2026.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d2e111">The accurate classification of infrasound signals is a cornerstone of global geophysical monitoring, essential for both natural hazard early warning systems (e.g., volcanic eruptions, debris flows, and earthquakes) and the verification of the Comprehensive Nuclear-Test-Ban Treaty. However, the development of reliable automated systems is hindered by the inherent scarcity of representative event data, particularly for rare extreme events, as well as the presence of complex, non-stationary background interferences. To maximize the diagnostic value of limited geophysical datasets, this paper proposes Infra-Net, a novel parallel decision-making network driven by multi-view feature learning and a confidence-based fusion mechanism. We introduce a logarithmic wavelet scattering transform to produce robust, mathematically grounded feature representations. Unlike conventional methods that process scattering matrices holistically, our approach treats individual columns as independent feature vectors, providing multiple localized perspectives of the same acoustic event. Architecturally, Infra-Net utilizes a dual-branch structure to simultaneously capture multi-scale spatial features and discriminative temporal patterns. These parallel evaluations are synthesized through a custom confidence-based fusion module, which employs weighted averaging and an inner-product mechanism to ensure a stable and comprehensive final classification. Tested on both public infrasound datasets and real-world Comprehensive Nuclear-Test-Ban Treaty Organization (CTBTO)-measured data, Infra-Net achieved accuracies of 100 % and 82.07 %, respectively. These results demonstrate that Infra-Net offers a highly robust soft computing solution for enhancing Earth system monitoring and the reliable identification of natural hazards amidst anthropogenic noise.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d2e123">Natural hazards such as earthquakes, volcanic eruptions, and debris flows pose significant threats to human society and ecosystems. To improve the timeliness and accuracy of early warning systems, it is essential to conduct multi-dimensional, real-time monitoring of these geodynamic processes. Among various monitoring technologies, infrasound sensing has gained prominence due to its unique physical properties. Infrasound refers to acoustic waves with frequencies below the lower threshold of human hearing, typically defined as being under 20 Hz (Bryan et al., 2018). These waves are characterized by low atmospheric absorption and the ability to propagate over vast distances, making them ideal for detecting remote geophysical events. Such waves are widely generated by natural phenomena, including ground vibrations from earthquakes (Alegria et al., 2015; Li et al., 2016), pressure releases from volcanic eruptions (Liu et al., 2014; Toney et al., 2022), and the geomorphic mass movements associated with debris flows (Kogelnig et al., 2014; Marchetti et al., 2019). Consequently, infrasound signals serve as powerful tools for sensing distant disasters and as vital carriers of information regarding energy exchange within the Earth system. With the collaborative development of sensor networks and signal processing techniques, the use of automated algorithms to classify and recognize infrasound signals has become central to enhancing hazard monitoring effectiveness. However, the accurate identification of these signals in practical applications faces significant obstacles. First, while human activities such as rocket launches (Dai et al., 2021; Wen et al., 2019) and chemical explosions (Koch and Pilger, 2018; Park et al., 2018) generate similar acoustic signatures, disaster mitigation tasks require the precise discrimination of natural hazards from these anthropogenic sources to avoid false alarms. Second, because major natural hazards, particularly extreme events, are sporadic and rare, it is often difficult for researchers to obtain sufficient field observation samples. This “data-scarce” dilemma severely constrains the generalization capabilities of traditional deep learning models, leading to suboptimal performance when encountering complex environmental interferences.</p>
      <p id="d2e126">In recent years, the academic community has proposed several schemes for infrasound recognition. For instance, Thüring et al. (2015) utilized Fourier spectra and support vector machine (SVMs) to successfully reduce avalanche detection error rates from 65 % to 10 %. Albert and Linville (2020) achieved an average classification accuracy of only 74 % for volcanic activity and seismic events, and a notably lower overall recognition accuracy of 56 % for earthquakes, chemical explosions, mine explosions, and volcanic activity using convolutional neural networks (CNNs). Leng et al. (2022) compiled an infrasound dataset including debris flows and lightning, achieving an accuracy of 84.1 % using an enhanced LeNet-5 architecture. Pásztor et al. (2023) extracted time-domain, frequency-domain, and other features from infrasound signals generated by quarry blasts and industrial sources. By training SVM and random forest classifiers on these features, they attained <inline-formula><mml:math id="M1" display="inline"><mml:mrow><mml:mi>F</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> scores of 0.88 and 0.93. While these studies have advanced the field, limitations remain. Many models, such as the improved AlexNet proposed by Yuan et al. (2024), rely heavily on data augmentation techniques like slicing and translation to expand samples, which may not accurately reflect the underlying physical distribution of rare natural events. Furthermore, hybrid methods like the long short-term memory (LSTM)-prototypical network introduced by Zhao et al. (2024) require extensive computational iterations, which may limit their utility in time-sensitive monitoring scenarios where efficiency is paramount.</p>
      <p id="d2e139">In summary, the domain of infrasound signal classification is still characterized by four primary bottlenecks. First, deep learning models are often constrained by the scarcity of event data, where traditional augmentation fails to compensate for missing physical information. Second, the opaque nature of many neural networks complicates architecture tuning and necessitates more interpretable feature extraction methods. Third, single-stream architectures often fail to capture the multifaceted spatio-temporal dynamics inherent in non-stationary geophysical signals. Finally, systems lacking redundancy and consensus mechanisms are vulnerable to individual misclassifications, which can compromise the stability of the overall monitoring regime.</p>
      <p id="d2e142">To address these challenges, this study introduces Infra-Net, a parallel decision-making framework designed for robust event classification under data-scarce conditions. Instead of relying on conventional data augmentation, this work leverages multi-view feature learning and confidence-based fusion to maximize information gain from limited samples. The specific contributions of this work are as follows: <list list-type="bullet"><list-item>
      <p id="d2e147">Multi-view Feature Extraction: a wavelet scattering network (WSN) (Khan et al., 2017) is employed to extract mathematically grounded features. By treating each column of the scattering coefficients as an independent feature vector, the model evaluates the same acoustic event from multiple localized perspectives, thereby enhancing the description of signal non-stationarity without the need for artificial sample expansion.</p></list-item><list-item>
      <p id="d2e151">Feature Contrast Enhancement: a logarithmic transformation is applied to the scattering coefficients to amplify subtle variations between natural hazards and anthropogenic interference sources, forming a highly discriminative feature space.</p></list-item><list-item>
      <p id="d2e155">Parallel Decision Architecture: the Infra-Net architecture comprises two specialized branches, one focused on multi-scale spatial features and the other on critical temporal patterns.</p></list-item><list-item>
      <p id="d2e159">Confidence-Based Fusion: a novel decision-making module is introduced. By performing confidence-weighted averaging followed by a confidence-based inner product, the system dynamically integrates the parallel evaluations from both branches. This consensus-driven approach mitigates the impact of noise and enhances the overall robustness of the classification system.</p></list-item></list></p>
      <p id="d2e163">Extensive evaluations demonstrate that this model effectively addresses the challenges of classifying small-sample infrasound events. This advancement strengthens the capability to characterize critical geophysical phenomena, supporting both natural hazard early warning systems and the verification mission of the Comprehensive Nuclear-Test-Ban Treaty (CTBT) (Lee and Hong, 2024).</p>
      <p id="d2e166">The remainder of this paper is organized as follows. In the “Materials and Methods” section, we introduce the datasets used and the extraction process of logarithmic wavelet scattering features. The “Model Construction” section presents the proposed method and elaborates on its innovative aspects. The “Results” section details the experimental procedures and offers a comprehensive analysis of the outcomes. In the “Discussion” section, we compare the advantages of our proposed method with those of other existing methods and cast a vision for subsequent endeavors. The “Conclusion” section summarizes the principal findings of this study.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Materials and Methods</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Data sources and division</title>
      <p id="d2e185">The empirical foundation of this study consists of two distinct infrasound datasets that capture a wide array of geophysical phenomena and natural hazards. The first dataset is derived from the library of typical infrasonic signals (LOTIS), an open-source repository designed to facilitate the study of non-stationary atmospheric acoustic waves (Bryan et al., 2018; Zhao et al., 2024). These data were primarily recorded by the infrasound sensor array located at Windless Bight, Antarctica (sensor bandwidth: 0.01–0.2 Hz; the relative coordinates of example channels CHF1–CHF4 are shown in Fig. 1). This high-latitude monitoring site provides a unique environment for capturing a diverse range of Earth system events, including atmospheric gravity waves (AGW) triggered by auroras, microbaroms (MB), mountain associated waves (MAW), and infrasound generated by volcanic eruptions (VE). By including these natural hazard precursors alongside complex background noise, the dataset provides a rigorous baseline for testing the discriminative power of the proposed model. The dataset division follows the approach described in the reference, as detailed in Table 1.</p>

      <fig id="F1"><label>Figure 1</label><caption><p id="d2e190">Schematic diagram of the infrasound sensor array located at Windless Bight, Antarctica. The array consists of multiple infrasound sensors (channels CHF1–CHF4 are shown as examples, with a bandwidth of 0.01–0.2 Hz, and their relative Cartesian coordinates in kilometers: CHF1 at (0, 0), CHF2 at (<inline-formula><mml:math id="M2" display="inline"><mml:mo lspace="0mm">-</mml:mo></mml:math></inline-formula>2.4055, 5.6579), CHF3 at (5.4587, 3.0939), and CHF4 at (3.5853, <inline-formula><mml:math id="M3" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>1.0567)).</p></caption>
          <graphic xlink:href="https://nhess.copernicus.org/articles/26/4529/2026/nhess-26-4529-2026-f01.png"/>

        </fig>

<table-wrap id="T1"><label>Table 1</label><caption><p id="d2e216">The training set and test set division of LOTIS (Li et al., 2026).</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="5">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">AGW</oasis:entry>
         <oasis:entry colname="col3">MAW</oasis:entry>
         <oasis:entry colname="col4">MB</oasis:entry>
         <oasis:entry colname="col5">VE</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">training set</oasis:entry>
         <oasis:entry colname="col2">81</oasis:entry>
         <oasis:entry colname="col3">77</oasis:entry>
         <oasis:entry colname="col4">41</oasis:entry>
         <oasis:entry colname="col5">46</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Test set</oasis:entry>
         <oasis:entry colname="col2">34</oasis:entry>
         <oasis:entry colname="col3">33</oasis:entry>
         <oasis:entry colname="col4">17</oasis:entry>
         <oasis:entry colname="col5">19</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Total</oasis:entry>
         <oasis:entry colname="col2">115</oasis:entry>
         <oasis:entry colname="col3">110</oasis:entry>
         <oasis:entry colname="col4">58</oasis:entry>
         <oasis:entry colname="col5">65</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e313">The other infrasound dataset used in this study consists of field-measured data obtained from the Comprehensive Nuclear-Test-Ban Treaty Organization (CTBTO). Due to high frequency of occurrence and the close resemblance of waveform characteristics and time-frequency distributions to those of nuclear explosions, natural earthquakes and chemical explosions represent significant sources of interference (Wang et al., 2022; Chang and Cui, 2013). Therefore, this dataset specifically emphasizes infrasound events generated by natural earthquakes and chemical explosions. Table 2 lists the number of events and waveforms for each category of infrasound in the catalog. The monitoring system runs continuously and actively 24 h a day, 7 d a week. Because each infrasound event is simultaneously captured by multiple sensors within a single station array as well as by multiple stations across the network, a one-to-many relationship is established between an event and its recordings. As a result, every event produces multiple waveforms.</p>

<table-wrap id="T2"><label>Table 2</label><caption><p id="d2e319">Infrasound events and waveform numbers (Li et al., 2026).</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Event</oasis:entry>
         <oasis:entry colname="col2">Number of</oasis:entry>
         <oasis:entry colname="col3">Number of</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">events</oasis:entry>
         <oasis:entry colname="col3">waveforms</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Chemical explosions</oasis:entry>
         <oasis:entry colname="col2">28</oasis:entry>
         <oasis:entry colname="col3">386</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Earthquakes</oasis:entry>
         <oasis:entry colname="col2">8</oasis:entry>
         <oasis:entry colname="col3">403</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e386">Partial event information is summarized in Table 3.</p>

<table-wrap id="T3" specific-use="star"><label>Table 3</label><caption><p id="d2e392">Partial event information for chemical explosions and natural earthquakes (Li et al., 2026). The time is given in Universal Time (UT).</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="5">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="left"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Event time</oasis:entry>
         <oasis:entry colname="col2">Event trigger</oasis:entry>
         <oasis:entry colname="col3">Latitude</oasis:entry>
         <oasis:entry colname="col4">Longitude</oasis:entry>
         <oasis:entry colname="col5">Event</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">(yyyymmdd)</oasis:entry>
         <oasis:entry colname="col2">time</oasis:entry>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">20020920</oasis:entry>
         <oasis:entry colname="col2">00:38:03.0</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M4" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>31.0108</oasis:entry>
         <oasis:entry colname="col4">136.7756</oasis:entry>
         <oasis:entry colname="col5">Woomera Test Explosion</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">20050321</oasis:entry>
         <oasis:entry colname="col2">20:46:22.8</oasis:entry>
         <oasis:entry colname="col3">41.131</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M5" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>112.896</oasis:entry>
         <oasis:entry colname="col5">UTTR Missile Detonation</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">20050323</oasis:entry>
         <oasis:entry colname="col2">19:20:00.0</oasis:entry>
         <oasis:entry colname="col3">29.374</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M6" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>94.939</oasis:entry>
         <oasis:entry colname="col5">BP Refinery Explosion</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">20051113</oasis:entry>
         <oasis:entry colname="col2">05:41:00.0</oasis:entry>
         <oasis:entry colname="col3">43.85</oasis:entry>
         <oasis:entry colname="col4">126.57</oasis:entry>
         <oasis:entry colname="col5">Jilin Chemical Factory Explosion</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">20020104</oasis:entry>
         <oasis:entry colname="col2">19:38:08.0</oasis:entry>
         <oasis:entry colname="col3">32.4</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M7" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>115.1</oasis:entry>
         <oasis:entry colname="col5">Western US Earthquake</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">20020319</oasis:entry>
         <oasis:entry colname="col2">22:14:49.9</oasis:entry>
         <oasis:entry colname="col3">30.1</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M8" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>113.7</oasis:entry>
         <oasis:entry colname="col5">Gulf of California Earthquake</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">20041226</oasis:entry>
         <oasis:entry colname="col2">01:00:00.0</oasis:entry>
         <oasis:entry colname="col3">5</oasis:entry>
         <oasis:entry colname="col4">95</oasis:entry>
         <oasis:entry colname="col5">Sumatra <inline-formula><mml:math id="M9" display="inline"><mml:mi>M</mml:mi></mml:math></inline-formula> 9 Earthquake</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">20010904</oasis:entry>
         <oasis:entry colname="col2">12:45:53.0</oasis:entry>
         <oasis:entry colname="col3">37.1</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M10" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>104.7</oasis:entry>
         <oasis:entry colname="col5">Colorado Earthquake</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">20050410</oasis:entry>
         <oasis:entry colname="col2">11:15:00.0</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M11" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>1.7</oasis:entry>
         <oasis:entry colname="col4">99.8</oasis:entry>
         <oasis:entry colname="col5">Mentawai <inline-formula><mml:math id="M12" display="inline"><mml:mi>M</mml:mi></mml:math></inline-formula> 6.5 Earthquake</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e669">The amount of infrasound data used in this study is limited. To make more efficient use of the data, the infrasound data are divided into three sub-datasets, and it is ensured that infrasound data from the same event are included in only one subset, thereby preventing the risk of data leakage. Table 4 displays the number of samples and events in the three subsets.</p>

<table-wrap id="T4" specific-use="star"><label>Table 4</label><caption><p id="d2e676">Waveform quantity of each subset.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="5">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="left"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">Number of chemical</oasis:entry>
         <oasis:entry colname="col3">Contains partial chemical</oasis:entry>
         <oasis:entry colname="col4">Number of earthquake</oasis:entry>
         <oasis:entry colname="col5">Contains partial</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">explosion waveforms</oasis:entry>
         <oasis:entry colname="col3">explosion events</oasis:entry>
         <oasis:entry colname="col4">waveforms</oasis:entry>
         <oasis:entry colname="col5">earthquake events</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Subset 1</oasis:entry>
         <oasis:entry colname="col2">123</oasis:entry>
         <oasis:entry colname="col3">Woomera test explosion</oasis:entry>
         <oasis:entry colname="col4">121</oasis:entry>
         <oasis:entry colname="col5">Colorado</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">BP refinery explosion</oasis:entry>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5">western US</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Subset 2</oasis:entry>
         <oasis:entry colname="col2">144</oasis:entry>
         <oasis:entry colname="col3">UTTR missile detonation</oasis:entry>
         <oasis:entry colname="col4">154</oasis:entry>
         <oasis:entry colname="col5">Gulf of California</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Jilin chemical factory explosion</oasis:entry>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5">Sumatra</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Subset 3</oasis:entry>
         <oasis:entry colname="col2">119</oasis:entry>
         <oasis:entry colname="col3">Albania arms depot explosions</oasis:entry>
         <oasis:entry colname="col4">128</oasis:entry>
         <oasis:entry colname="col5">Mentawai</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <fig id="F2" specific-use="star"><label>Figure 2</label><caption><p id="d2e821">Schematic diagram of dataset division.</p></caption>
          <graphic xlink:href="https://nhess.copernicus.org/articles/26/4529/2026/nhess-26-4529-2026-f02.png"/>

        </fig>

      <p id="d2e830">The data division approach is designed to minimize the correlation between the test set and the training set as much as possible, ensuring the reliability of the classification outcomes (Tan et al., 2021). The division methodology is illustrated in Fig. 2, beginning with the division of data into three event-based subsets containing 244, 298, and 247 raw signals respectively. Each subset undergoes log-scattering feature extraction, generating six feature sets per original signal (when the number of scattering paths is six), resulting in a total of 4734 input groups. In the classification experiments, one feature subset was sequentially employed as the test set, while the other two were utilized for training. The final metric was derived by averaging the results from these three independent classification trials.</p>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Extraction of logarithmic wavelet scattering features</title>
<sec id="Ch1.S2.SS2.SSS1">
  <label>2.2.1</label><title>Wavelet scattering network</title>
      <p id="d2e848">Wavelet transform (Lilly and Olhede, 2010), as a method for analyzing non-stationary signals, is characterized by multi-resolution and scale variability. However, it lacks translation invariance (Li et al., 2019), which can lead to the loss of certain signal details during analysis. Specifically, when decomposing low-frequency signals, high-frequency components may be overlooked, resulting in the loss of detail in the high-frequency signal (Fan et al., 2022). In contrast, the wavelet scattering transform (WST) computes a semi-discrete wavelet transform of the signal (Souli et al., 2021) and then applies a nonlinear modulus operation. The resulting feature representation possesses desirable properties such as translation invariance and deformation stability, perfectly meeting the basic requirements for feature extractors in machine learning (Wiatowski and Bolcskei, 2018; Wang et al., 2018a, b; Huang et al., 2018; Liu et al., 2019).</p>

      <fig id="F3" specific-use="star"><label>Figure 3</label><caption><p id="d2e853">Schematic diagram of the wavelet scattering network.</p></caption>
            <graphic xlink:href="https://nhess.copernicus.org/articles/26/4529/2026/nhess-26-4529-2026-f03.png"/>

          </fig>

      <p id="d2e862">As a geometrically invariant feature extractor (Tuan, 2023; Buriro et al., 2021; Andén and Mallat, 2014), WST maintains robustness against translation, frequency shifts, and scale changes while preserving discriminative structures–crucial for non-stationary infrasound signals affected by atmospheric turbulence. The structure of the WSN is similar to that of a deep CNN and is primarily comprised of three components: signal input, feature extraction, and feature output. The feature extraction component is composed of multiple modules, each of which consists of wavelet convolution, nonlinearity, and averaging. The network's architecture is depicted in Fig. 3. Unlike CNNs, the filters in the WSN are predefined wavelet filters, whose parameters do not require learning through training samples (Ding et al., 2024), eliminating parameter training requirements and preventing overfitting in data-limited scenarios. In scenarios with small samples, WSNs typically achieve lower classification error rates compared to deep CNNs, thus offering certain advantages in the study of the classification for small-sample infrasound events.</p>
      <p id="d2e866">The WST achieves local translation-invariant descriptors of infrasound signals through time-domain averaging

              <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M13" display="block"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>F</mml:mi><mml:mo>*</mml:mo><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            where <inline-formula><mml:math id="M14" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the scattering coefficient of <inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. This descriptor eliminates all of the high-frequency information from the infrasound signal, thus extracting its low-frequency content while maintaining translation invariance. The high-frequency information is obtained by utilizing the modulus operation of the wavelet transform.

              <disp-formula id="Ch1.E2" content-type="numbered"><label>2</label><mml:math id="M16" display="block"><mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mfenced close="|" open="|"><mml:mrow><mml:mi>F</mml:mi><mml:mo>*</mml:mo><mml:msub><mml:mi mathvariant="italic">ψ</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            where <inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the scalogram coefficient of <inline-formula><mml:math id="M18" display="inline"><mml:mrow><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M19" display="inline"><mml:mi mathvariant="italic">ψ</mml:mi></mml:math></inline-formula> is the high-frequency wavelet function, and <inline-formula><mml:math id="M20" display="inline"><mml:mi mathvariant="italic">ϕ</mml:mi></mml:math></inline-formula> is the low-pass filter. Equation (2) represents the high-frequency information at scale <inline-formula><mml:math id="M21" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> and acquires deformation stability through the modulus operation of the nonlinear wavelet transform (Bruna and Mallat, 2011, 2013).</p>
      <p id="d2e1033">When Eq. (2) is used as the input for the first-order scattering transform, the following equations are obtained (Tan et al., 2024; Liu et al., 2025; Shirodkar et al., 2025):

                  <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M22" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E3"><mml:mtd><mml:mtext>3</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>S</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mfenced open="|" close="|"><mml:mrow><mml:mi>F</mml:mi><mml:mo>*</mml:mo><mml:msub><mml:mi mathvariant="italic">ψ</mml:mi><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mo>*</mml:mo><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E4"><mml:mtd><mml:mtext>4</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>U</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mfenced close="|" open="|"><mml:mrow><mml:mfenced close="|" open="|"><mml:mrow><mml:mi>F</mml:mi><mml:mo>*</mml:mo><mml:msub><mml:mi mathvariant="italic">ψ</mml:mi><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mo>*</mml:mo><mml:msub><mml:mi mathvariant="italic">ψ</mml:mi><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            By iteratively applying the aforementioned process, scattering coefficients of any order can be obtained. For any <inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mo>≥</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>, the iteration of the wavelet modulus convolution is

              <disp-formula id="Ch1.E5" content-type="numbered"><label>5</label><mml:math id="M24" display="block"><mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mfenced close="|" open="|"><mml:mrow><mml:mfenced close="|" open="|"><mml:mrow><mml:mfenced open="|" close="|"><mml:mrow><mml:mi>F</mml:mi><mml:mo>*</mml:mo><mml:msub><mml:mi mathvariant="italic">ψ</mml:mi><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mo>*</mml:mo><mml:mi mathvariant="normal">…</mml:mi></mml:mrow></mml:mfenced><mml:mo>*</mml:mo><mml:msub><mml:mi mathvariant="italic">ψ</mml:mi><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>

            The <inline-formula><mml:math id="M25" display="inline"><mml:mi>m</mml:mi></mml:math></inline-formula>-order scattering coefficient is

              <disp-formula id="Ch1.E6" content-type="numbered"><label>6</label><mml:math id="M26" display="block"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mfenced close="|" open="|"><mml:mrow><mml:mfenced open="|" close="|"><mml:mrow><mml:mfenced close="|" open="|"><mml:mrow><mml:mi>F</mml:mi><mml:mo>*</mml:mo><mml:msub><mml:mi mathvariant="italic">ψ</mml:mi><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mo>*</mml:mo><mml:mi mathvariant="normal">…</mml:mi></mml:mrow></mml:mfenced><mml:mo>*</mml:mo><mml:msub><mml:mi mathvariant="italic">ψ</mml:mi><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo>*</mml:mo><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>

            The final feature vector is composed of the scattering coefficients of each order.</p>
</sec>
<sec id="Ch1.S2.SS2.SSS2">
  <label>2.2.2</label><title>Extraction of logarithmic wavelet scattering features</title>
      <p id="d2e1370">The spectral distribution of a signal determines the numerical distribution of its wavelet scattering features. If a signal contains abundant high-frequency components, the scattering features generated in the high-frequency region of the wavelet transform may have larger numerical values. The focus of this research is on infrasound signals, which are characterized by relatively low frequencies. Analysis reveals that the frequency distribution of the two types of infrasound event signals under investigation typically ranges from a few tenths to several hertz. Consequently, the wavelet scattering features produced also exhibit relatively small magnitudes. The logarithmic function, being monotonic within its domain, preserves the relative relationships between data points. Moreover, due to the properties of the logarithmic function, it is more sensitive to differences in smaller values compared to differences in larger values, which can further highlight the disparities between features. Therefore, logarithmic wavelet scattering features are utilized as the inputs of the classifier, which enhances the local characteristics of the signal and makes subtle variations more pronounced, thereby improving the expressive power of the features. As shown in Fig. 4, which depicts the wavelet scattering feature extraction framework, the extraction of logarithmic wavelet scattering features can be performed by utilizing the following steps. Following Eqs. (1) to (6), third-order wavelet scattering features are extracted. Thereafter, the logarithm of the extracted wavelet scattering features is computed to obtain the logarithmic wavelet scattering features. If higher-order logarithmic wavelet scattering coefficients are required, Eqs. (5) and (6) can be repeated. However, in practical applications, as the number of layers increases, the energy of the infrasound signal will continuously decrease (Lu et al., 2023), leading to a significant decrease in the discriminability of the features obtained. Therefore, the number of scattering layers is set to three (Zhang et al., 2024b).</p>

      <fig id="F4" specific-use="star"><label>Figure 4</label><caption><p id="d2e1375">Schematic diagram of the wavelet scattering feature extraction framework.</p></caption>
            <graphic xlink:href="https://nhess.copernicus.org/articles/26/4529/2026/nhess-26-4529-2026-f04.png"/>

          </fig>

      <p id="d2e1384">The input infrasound signal has a sampling frequency of 20 Hz and a time-invariant scale size of 150 s when an analytic Morlet wavelet is employed (Stéphane, 2009). The number of wavelet filters per octave is set to [8 2 1] (Al and Khushaba, 2023). Taking a measured infrasound signal as an example, which has a data size of 1 <inline-formula><mml:math id="M27" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 3000, is input into the wavelet scattering network to extract the wavelet scattering features, resulting in a data size of 307 <inline-formula><mml:math id="M28" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 6, meaning each wavelet scattering feature consists of 6 columns of wavelet scattering coefficients. Each column corresponds to specific wavelet scales and scattering paths.</p>

      <fig id="F5" specific-use="star"><label>Figure 5</label><caption><p id="d2e1404">Three-layer wavelet scattering features comparing a chemical explosion (left) with an earthquake (right). Key discriminative time windows (60–100 s in Layer 1, after 120 s in Layer 3) and frequency-dependent energy separations (Layer 2) are indicated. This multi-layer visualization clarifies how scattering features capture source-specific signatures, supporting robust event classification.</p></caption>
            <graphic xlink:href="https://nhess.copernicus.org/articles/26/4529/2026/nhess-26-4529-2026-f05.jpg"/>

          </fig>

      <p id="d2e1413">Unlike conventional approaches that concatenate the six columns of scattering features into a single, high-dimensional input matrix (Lone and Aydin, 2023; Priya et al., 2021; Singhal and Kumar, 2024), this study processes each column of scattering coefficients as an independent feature vector fed into a parallel evaluation framework. This strategy is grounded in the theoretical foundation of the wavelet scattering transform: scattering coefficients at different levels capture rich, non-stationary information from multi-scale and multi-resolution perspectives. Each column of scattering features possesses a distinct mathematical expression and physical significance (as shown in Eqs. 1 to 6), representing specific, complementary attributes of the same acoustic event. By treating these columns as distinct inputs for parallel analysis, the proposed architecture comprehensively preserves the multi-level information of the original signal without the risk of feature redundancy or inter-channel interference that often occurs in flattened inputs. More importantly, this processing strategy enables the generation of multiple independent confidence representations for a single signal. This multi-view characterization lays the essential foundation for the subsequent confidence-based decision-making mechanism, significantly enhancing the system's robustness against complex background interferences. As illustrated in Fig. 5, the visualization of the first three layers of wavelet scattering features from two representative signal samples clearly demonstrates that features at different levels provide complementary and discriminative insights, which are crucial for distinguishing nuclear events from natural phenomena.</p>
      <p id="d2e1416">As shown in Fig. 5, in the first layer, the characteristic frequency of chemical explosions is generally slightly lower than those of earthquakes, consistent with the more concentrated energy release typical of explosive sources. A clear difference in signal energy appears between 60 and 100 s, reflecting the distinct source mechanisms: chemical explosions typically produce impulsive, short-duration signals while earthquakes generate more extended waveforms from complex fault ruptures. Outside this time window, the signal energy distributions show substantial similarity due to common propagation effects.</p>
      <p id="d2e1419">In the second layer, the range of distinct signal energy distribution expands significantly. From 60 s onward, marked differences emerge in the spectral energy distribution, where explosion signals demonstrate energy concentrations tending toward lower frequencies compared to earthquakes. This pattern aligns with the fundamental physical distinction between the relatively simple explosive source mechanism and the more complex double-couple mechanism of earthquakes. However, overall discrimination remains challenging in the first two layers due to limited frequency resolution.</p>
      <p id="d2e1422">The third layer proves crucial as its enhanced frequency resolution reveals pronounced differences in signal energies after 120 s. This improved resolution effectively captures the more rapid amplitude decay characteristic of explosion signals compared to the sustained energy release of earthquake events, thereby facilitating the extraction of subtle discriminative features.</p>
      <p id="d2e1425">The strategic utilization of scattering features from all three layers as parallel inputs enables a comprehensive analysis of these heterogeneous characteristics. Rather than increasing the sample count, this multi-vector approach enhances the granularity of event representation, which is particularly valuable in small-sample-size scenarios where every bit of discriminative information is critical. The layer-specific information content confirms that evaluating scattering features from different layers as distinct feature views effectively captures the full diversity of the signal's physical properties. These results provide a more comprehensive and reliable diagnostic framework for model training, which ultimately strengthens the precision and stability of automated systems dedicated to the real-time detection of natural hazards.</p>
</sec>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Model Construction</title>
      <p id="d2e1438">Infrasound signals exhibit significant characteristics including long-term dependence, multi-scale features, and non-stationarity. Leveraging these properties, a classification model named Infra-Net is designed, comprising two branch networks and a confidence-based decision-making module to enhance classification accuracy and robustness, as shown in Fig. 6. The Infra-Net consists of five components: input, Branch 1 network, Branch 2 network, confidence-based decision-making module, and output. The design objectives of the principal modules are detailed as follows.</p>

      <fig id="F6" specific-use="star"><label>Figure 6</label><caption><p id="d2e1443">Schematic diagram of Infra-Net.</p></caption>
        <graphic xlink:href="https://nhess.copernicus.org/articles/26/4529/2026/nhess-26-4529-2026-f06.png"/>

      </fig>

<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Branch 1: GA-BiGRU</title>
      <p id="d2e1459">For Branch 1, the scattering features derived from the wavelet scattering transform exhibit strict temporal locality preservation, necessitating a sequential modeling network that balances the capture of long-range dependencies with computational efficiency. Unlike the BiLSTM (Bidirectional Long Short-Term Memory), the BiGRU (Bidirectional Gated Recurrent Unit) utilizes a simplified gating mechanism, which significantly reduces parameter complexity while maintaining the ability to model long-term dependencies in sequential data (Cho et al., 2014; Hochreiter and Schmidhuber, 1997). This design enhances parameter convergence efficiency on limited-sample infrasound datasets and effectively mitigates the risk of overfitting.</p>
      <p id="d2e1462">However, the traditional BiGRU architecture utilizes only the final hidden state as the sequential representation, while the intermediate hidden states generated during the process are not directly incorporated into the final output computation. This approach implicitly assumes the Markov property, where <inline-formula><mml:math id="M29" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">h</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is expected to encode the entire historical information. According to the principle of entropy increase in information theory, as the sequence length grows, the information entropy of the hidden state exhibits an exponential decay trend, leading to the gradual loss of early temporal features. Consequently, the final hidden state overly focuses on the final features, affecting the model's ability to model the entire sequence. To address this issue, a Global Attention (GA) mechanism (Luong et al., 2015; Zhang et al., 2024a) is introduced, which aims to mitigate information loss and enhance global interactive representation. This mechanism preserves information across both channel and spatial dimensions, thereby improving the performance of the neural network. The schematic diagram of the GA-BiGRU architecture is shown in Fig. 7.</p>

      <fig id="F7"><label>Figure 7</label><caption><p id="d2e1478">Schematic diagram of the GA-BiGRU.</p></caption>
          <graphic xlink:href="https://nhess.copernicus.org/articles/26/4529/2026/nhess-26-4529-2026-f07.png"/>

        </fig>

      <p id="d2e1488">The scalar <inline-formula><mml:math id="M30" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents the weight of the hidden state <inline-formula><mml:math id="M31" display="inline"><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> at time step <inline-formula><mml:math id="M32" display="inline"><mml:mi>s</mml:mi></mml:math></inline-formula> relative to all hidden states.

            <disp-formula id="Ch1.E7" content-type="numbered"><label>7</label><mml:math id="M33" display="block"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>exp⁡</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">score</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:msup><mml:mi>s</mml:mi><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:munder><mml:mi>exp⁡</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="normal">score</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:msup><mml:mi>s</mml:mi><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:msub><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:math></disp-formula>

          By combining the weights <inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> of all <inline-formula><mml:math id="M35" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> hidden states, the weight vector <inline-formula><mml:math id="M36" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is obtained. The term <inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">c</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes the context vector at time step <inline-formula><mml:math id="M38" display="inline"><mml:mi>s</mml:mi></mml:math></inline-formula>, which is computed by multiplying the hidden state <inline-formula><mml:math id="M39" display="inline"><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> with its corresponding weight <inline-formula><mml:math id="M40" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. The global context vector  <inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">c</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> for the entire sample is computed by summing   across all time steps.

            <disp-formula id="Ch1.E8" content-type="numbered"><label>8</label><mml:math id="M42" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">c</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:msub><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>)</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the number of sequence length. The target hidden state <inline-formula><mml:math id="M44" display="inline"><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>  and context vector <inline-formula><mml:math id="M45" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">c</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are concatenated to form an attention-enhanced vector <inline-formula><mml:math id="M46" display="inline"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">h</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover></mml:math></inline-formula>.

            <disp-formula id="Ch1.E9" content-type="numbered"><label>9</label><mml:math id="M47" display="block"><mml:mrow><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">h</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mi mathvariant="normal">tanh</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">W</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>[</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>;</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>

          Finally, <inline-formula><mml:math id="M48" display="inline"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">h</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover></mml:math></inline-formula> is fed into a Softmax layer to produce the classification probability <inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.

            <disp-formula id="Ch1.E10" content-type="numbered"><label>10</label><mml:math id="M50" display="block"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:mi>y</mml:mi><mml:mo>&lt;</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="normal">softmax</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">W</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">h</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>

          Through the aforementioned method, GA learns the weights of each time step of the BiGRU, comprehensively considering all hidden states and assigning scores to emphasize important features while suppressing irrelevant ones. This method not only fully leverages the strengths of BiGRU in processing sequential data but also addresses the limitation of traditional methods in insufficiently capturing key features through the attention mechanism, thereby significantly enhancing the model's performance.</p>
</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Branch 2: Multi-Scale Convolutional Infrasound Network (MSCI-Net)</title>
      <p id="d2e1933">The design of Branch 2 aims to leverage the distinct time-frequency feature distributions between the two types of infrasound events (Fig. 5). To this end, this branch incorporates a lightweight CNN architecture, centered around a multi-scale convolutional structure and a hierarchical feature extraction mechanism. Specifically, this branch utilizes parallel multi-scale convolutional kernels to capture different characteristics across varying receptive fields: smaller-scale kernels focus on capturing fine-grained fluctuations and local structures in the signal, while larger-scale kernels are used to extract energy distribution trends and macro patterns over extended time ranges. Through this multi-scale collaborative mechanism, the model can simultaneously respond to both local anomalies and global evolutionary patterns in infrasound signals. To further address the risk of overfitting in limited-data scenarios, this branch abandons traditional deep convolutional networks in favor of a shallower architecture with fewer parameters and a more streamlined structure. This design enhances its sensitivity and robustness to key spatial features (such as energy distribution and structural patterns). Through the aforementioned multi-scale and hierarchical processing, Branch 2 can effectively capture discriminative spatial structures across different time intervals, significantly improving the differentiation capability of infrasound events. The architecture is shown in Fig. 8.</p>

      <fig id="F8" specific-use="star"><label>Figure 8</label><caption><p id="d2e1938">Schematic diagram of the MSCI-Net.</p></caption>
          <graphic xlink:href="https://nhess.copernicus.org/articles/26/4529/2026/nhess-26-4529-2026-f08.png"/>

        </fig>

      <p id="d2e1947">Four convolution layers are utilized to extract features, ensuring the lightweight nature of the MSCI-Net. The first two layers utilize kernels of sizes 9 <inline-formula><mml:math id="M51" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 1 and 7 <inline-formula><mml:math id="M52" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 1 to expand the receptive field, enabling the rapid capture of global features, which are essential for understanding the overall structure and trends of infrasound signals. In the third layer, multiple convolution layers are parallelized with varying kernel sizes at the same network level (Li et al., 2024). This allows for the fusion of information filtered at different scales and effectively reduces information loss without increasing the network depth. Before the fully connected layer, a 1 <inline-formula><mml:math id="M53" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 1 convolutional kernel is applied. This is specifically designed for small-sample infrasound signal processing, aiming to reduce parameter counts and enhance nonlinearity, thereby improving the modeling capability for infrasound signals. The LeakyReLU activation function is introduced after the convolution layers to replace the conventional ReLU (Rectified Linear Unit). LeakyReLU retains a small gradient in the negative region, mitigating information loss and gradient vanishing issues caused by feature compression, as detailed in Eq. (11).

            <disp-formula id="Ch1.E11" content-type="numbered"><label>11</label><mml:math id="M54" display="block"><mml:mrow><mml:mi mathvariant="italic">σ</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mfenced open="{" close=""><mml:mtable class="array" columnalign="left left"><mml:mtr><mml:mtd><mml:mi>x</mml:mi></mml:mtd><mml:mtd><mml:mrow><mml:mi mathvariant="normal">if</mml:mi><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mi>x</mml:mi><mml:mo>≥</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">γ</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mi mathvariant="normal">if</mml:mi><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mi>x</mml:mi><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mfenced><mml:mo>(</mml:mo><mml:mi mathvariant="italic">γ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.01</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>

          Additionally, a small stride is adopted, allowing the convolutional kernels to scan the signal step-by-step, thereby capturing both local details and global features. To enhance translational invariance and improve the MSCI-Net's robustness to minor temporal shifts, MaxPooling layers are incorporated. Furthermore, to address the non-stationarity of infrasound signals and their sensitivity to environmental noise, which complicates the data distribution, a batch normalization layer is employed to accelerate convergence, stabilize the training process, and suppress noise interference. Dropout is also introduced to randomly deactivate a portion of neurons (i.e., the fundamental computational units that receive input signals, compute a weighted sum, and apply a nonlinear activation function) during training, reducing MSCI-Net complexity and preventing overfitting. Ultimately, the MSCI-Net employs two fully connected layers that serve as a critical bridge from feature integration to classification decision-making. While convolutional layers excel at extracting local features, these features remain spatially or temporally fragmented. The first fully connected layer (with 64 neurons) performs global integration of these highly abstract yet scattered “local feature fragments.” Through weighted summation and nonlinear activation functions, it learns complex high-order combinatorial relationships among them, thereby forming a compact global representation that incorporates all contextual information. Subsequently, the second fully connected layer maps this 64-dimensional global representation to the final sample label space. Each of neurons directly corresponds to an output category. Subsequently, the Softmax function generates the classification probability for each branch. This design, through multi-scale convolutional kernel configurations, a lightweight architecture, and anti-interference mechanisms, significantly enhances the extraction capability for complex infrasound signal features and the robustness of classification, while maintaining computational efficiency.</p>
      <p id="d2e2033">Through the strategic use of multi-scale convolutional kernels, a lightweight architecture, and robust anti-interference mechanisms, this branch significantly boosts the Infra-Net's ability to extract complex features from infrasound signals and enhances classification robustness.</p>
</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>Confidence-based decision-making module</title>
      <p id="d2e2044">Leveraging information fusion theory (Zhang et al., 2025, 2024c; Dai et al., 2024), Infra-Net employs its novel confidence-based decision-making module to integrate the results of the two branches, achieving enhanced classification robustness.</p>
      <p id="d2e2047">As shown in Fig. 9, the confidence-based decision-making module is conducted in the following stages. In the information collection stage, after feature extraction through the logarithmic wavelet scattering network, a single 1 <inline-formula><mml:math id="M55" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M56" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> infrasound signal yields a <inline-formula><mml:math id="M57" display="inline"><mml:mrow><mml:mi>q</mml:mi><mml:mo>×</mml:mo><mml:mi>l</mml:mi></mml:mrow></mml:math></inline-formula> set of logarithmic wavelet scattering features. Each <inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:mi>q</mml:mi><mml:mo>×</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> scattering feature is treated as a distinct feature view for parallel assessment. Subsequently, the logarithmic wavelet scattering features are input into both the MSCI-Net and GA-BiGRU networks, which independently produce <inline-formula><mml:math id="M59" display="inline"><mml:mi>l</mml:mi></mml:math></inline-formula> sets of classification probability values corresponding to the same signal, denoted as <inline-formula><mml:math id="M60" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>B</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M61" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, respectively, where (<inline-formula><mml:math id="M62" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>, 2, …, <inline-formula><mml:math id="M63" display="inline"><mml:mi>l</mml:mi></mml:math></inline-formula>), <inline-formula><mml:math id="M64" display="inline"><mml:mi>q</mml:mi></mml:math></inline-formula> denotes the dimension of the scattering coefficients and <inline-formula><mml:math id="M65" display="inline"><mml:mi>l</mml:mi></mml:math></inline-formula> represents the number of scattering paths.</p>

      <fig id="F9"><label>Figure 9</label><caption><p id="d2e2187">Schematic diagram of information fusion structure.</p></caption>
          <graphic xlink:href="https://nhess.copernicus.org/articles/26/4529/2026/nhess-26-4529-2026-f09.png"/>

        </fig>

      <p id="d2e2197">At the local fusion node, for each signal, the average of the <inline-formula><mml:math id="M66" display="inline"><mml:mi>l</mml:mi></mml:math></inline-formula> sets of probability values output by each network are  calculated, which can be expressed as follows:

                <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M67" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E12"><mml:mtd><mml:mtext>12</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>A</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>B</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mfenced close=")" open="("><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>l</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>l</mml:mi></mml:munderover><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>l</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>l</mml:mi></mml:munderover><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>B</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E13"><mml:mtd><mml:mtext>13</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mfenced close=")" open="("><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>l</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>l</mml:mi></mml:munderover><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>l</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>l</mml:mi></mml:munderover><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula></p>
      <p id="d2e2379">The <inline-formula><mml:math id="M68" display="inline"><mml:mi>l</mml:mi></mml:math></inline-formula> collection nodes per branch are merged into one using the confidence-based averaging method (Eqs. 12–13), resulting in local integrated classification probabilities being obtained for each branch. This approach calculates the decision probability for a single signal across two branches by averaging the confidence of each column of scattering coefficients from the same signal, thereby mitigating the potential bias of a single-view representation through ensemble evaluation.</p>
      <p id="d2e2390">At the global fusion output node, to achieve a complementary advantage between the two branches, the second step of the confidence-based decision-making module is necessary. The final event category confidence is defined as the inner product of the output probabilities from both networks, which is denoted as

            <disp-formula id="Ch1.E14" content-type="numbered"><label>14</label><mml:math id="M69" display="block"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>A</mml:mi></mml:msub><mml:mo>⋅</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>B</mml:mi></mml:msub><mml:mo>⋅</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>

          The ultimate classification outcome is determined by the event category corresponding to <inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:mi mathvariant="normal">Max</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. The inner product operation in Eq. (14) combines the confidence levels of two networks through multiplication, emphasizing the joint support of both branches for the same category. The magnitude of the inner product value directly reflects the strength of the joint support of the two branches for that category. Therefore, the inner product value can serve as a measure of joint confidence, which helps enhance the reliability of the final decision.</p>
      <p id="d2e2474">In summary, by integrating temporal, spatial, and logarithmic wavelet scattering features, Infra-Net constructs a more diversified feature space. Subsequently, through a parallel network architecture and a confidence fusion strategy, it effectively enhances the reliability of the results, achieving robust and accurate classification of interference infrasound signals in natural hazards monitoring.</p>
</sec>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Results</title>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>Experimental parameter settings</title>
      <p id="d2e2493">Before training the Infra-Net, it is necessary to customize the Infra-Net's parameters. The number of hidden layer units in the BiGRU, the learning rate for training, and the maximum number of iterations all impact the recognition results. The optimal combination of these three parameters is determined using the enumeration method and grid search optimization (Abbaszadeh et al., 2022). The BiGRU network includes two layers, which contain 512 and 256 hidden layer units. The Adam optimizer is employed, and cross-entropy is utilized as the loss function. The initial learning rate is 0.001, it is reduced to 0.0001 after 400 training epochs, and the maximum number of epochs is set to 500, with a minimum of 400 training samples per epoch.</p>
</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>Experimental evaluation index</title>
      <p id="d2e2505">As shown in Table 5, the confusion matrix allows us to compute four statistical quantities: true positives, false positives, false negatives, and true negatives.</p>

<table-wrap id="T5"><label>Table 5</label><caption><p id="d2e2511">Confusion matrix (Li et al., 2026).</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Actual/Forecast</oasis:entry>
         <oasis:entry colname="col2">Positive</oasis:entry>
         <oasis:entry colname="col3">Negative</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Positive</oasis:entry>
         <oasis:entry colname="col2">True positive (TP)</oasis:entry>
         <oasis:entry colname="col3">False negative (FN)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Negative</oasis:entry>
         <oasis:entry colname="col2">False positive (FP)</oasis:entry>
         <oasis:entry colname="col3">True negative (TN)</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e2567">Three indices are utilized to evaluate the classification performance of the Infra-Net: the accuracy (ACC), precision (<inline-formula><mml:math id="M71" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula>), and <inline-formula><mml:math id="M72" display="inline"><mml:mrow><mml:mi>F</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> score. The recall (<inline-formula><mml:math id="M73" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula>), the rate of true positive identification, is also considered. The specific definitions of each index are as follows:

                <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M74" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E15"><mml:mtd><mml:mtext>15</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mi mathvariant="normal">ACC</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi mathvariant="normal">TP</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">TN</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">TP</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">FP</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">TN</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">FN</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E16"><mml:mtd><mml:mtext>16</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mi mathvariant="normal">TP</mml:mi><mml:mrow><mml:mi mathvariant="normal">TP</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">FP</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E17"><mml:mtd><mml:mtext>17</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mi mathvariant="normal">TP</mml:mi><mml:mrow><mml:mi mathvariant="normal">TP</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">FN</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E18"><mml:mtd><mml:mtext>18</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mi>F</mml:mi><mml:mn mathvariant="normal">1</mml:mn><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>×</mml:mo><mml:mi>R</mml:mi><mml:mo>×</mml:mo><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>R</mml:mi><mml:mo>+</mml:mo><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

          In addition, due to the highly pronounced uneven distribution of samples across various categories in the LOTIS dataset, Cohen's Kappa was further employed as a supplementary metric for evaluating model performance, to prevent the misleadingly optimistic assessment produced by traditional metrics such as accuracy under imbalanced data conditions. Cohen's Kappa is an important statistical measure used to evaluate the consistency and reliability between the predictions of a classification model and the true labels. Unlike accuracy, which only focuses on the proportion of correctly classified samples, this coefficient incorporates a correction for “random agreement”, eliminating the influence of random factors in the classification results. Thus, it provides a more robust and accurate measure of model performance. Especially in scenarios with highly skewed class distributions, Cohen's Kappa better reflects the model's true discriminative capability beyond random guessing. Its calculation formula is as follows:

            <disp-formula id="Ch1.E19" content-type="numbered"><label>19</label><mml:math id="M75" display="block"><mml:mrow><mml:mtext>Cohen's Kappa</mml:mtext><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where, <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the observed proportional agreement, representing the actual proportion of agreement between evaluators across all samples. <inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the expected random agreement proportion, calculated assuming evaluators classify independently and randomly.</p>
</sec>
<sec id="Ch1.S4.SS3">
  <label>4.3</label><title>Experiment and result analysis</title>
      <p id="d2e2789">First, experiments were conducted using the LOTIS dataset. Figure 10 shows classification results for Branch 1 (Fig. 10a), Branch 2 (Fig. 10b), and Infra-Net (Fig. 10c) using logarithmic wavelet scattering features. The limited data length only permits a threefold scattering feature path expansion. Experiments (a) and (b) generate three independent classifications per signal without decision fusion, yielding input counts of original signals <inline-formula><mml:math id="M78" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 3. Conversely, (c)'s confidence-based decision-making module fuses three results into single outputs, maintaining original signal counts. Thus, (a) and (b) contain threefold more features than (c). It can be seen that the classification accuracy reached above 98 % when using only one of the branches designed in this study, which is comparable to or even better than the classification results reported by Bryan et al. (2018) and Zhao et al. (2024) on the same dataset, both in terms of accuracy and time efficiency. Notably, compared with Zhao et al., which requires 10 000 iterations to achieve a maximum accuracy of 99.07 %, Branch 2 only needs 300 iterations to reach a result of 99.35 %, and the classification accuracy for AGW, MAW and VE events can all reach 100 %. Ultimately, when using Infra-Net for classification, the overall classification accuracy stabilizes at 100 %, yielding very excellent results. When facing the classification of different categories of infrasound signals, it is only necessary to adjust the number of neurons in the final fully connected layer of Infra-Net to match the corresponding number of categories, without requiring further modifications to the model architecture.</p>

      <fig id="F10"><label>Figure 10</label><caption><p id="d2e2801">Comparison of classification results for four types of infrasound events using three network configurations: <bold>(a)</bold> Branch 1, <bold>(b)</bold> Branch 2, and <bold>(c)</bold> Infra-Net.</p></caption>
          <graphic xlink:href="https://nhess.copernicus.org/articles/26/4529/2026/nhess-26-4529-2026-f10.png"/>

        </fig>

      <p id="d2e2819">For the comparative experiments, state-of-the-art wavelet scattering-based methods along with advanced techniques currently employed in infrasound classification were selected, including WST-AMResNet-18 (Liu et al., 2025), WST-Trans (Shirodkar et al., 2025), WST-BiLSTM (Zhang et al., 2024b), and MS-SE-ResNet (Tan et al., 2024). The results are summarized in Table 6.</p>

<table-wrap id="T6"><label>Table 6</label><caption><p id="d2e2826">The training set and test set division of LOTIS.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">ACC</oasis:entry>
         <oasis:entry colname="col3">Cohen's Kappa</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">(%)</oasis:entry>
         <oasis:entry colname="col3">(%)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">MS-SE-ResNet</oasis:entry>
         <oasis:entry colname="col2">96.12</oasis:entry>
         <oasis:entry colname="col3">94.65</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">WST-Trans</oasis:entry>
         <oasis:entry colname="col2">98.06</oasis:entry>
         <oasis:entry colname="col3">96.00</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">WST-BiLSTM</oasis:entry>
         <oasis:entry colname="col2">98.06</oasis:entry>
         <oasis:entry colname="col3">97.32</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">WST-AMResNet-18</oasis:entry>
         <oasis:entry colname="col2">94.17</oasis:entry>
         <oasis:entry colname="col3">91.96</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Infra-Net</oasis:entry>
         <oasis:entry colname="col2">100</oasis:entry>
         <oasis:entry colname="col3">100</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e2928">As can be observed from the Table 6, Infra-Net not only achieved the highest classification accuracy but also exhibited a significantly higher Cohen's Kappa compared to the other three models. The elevated Kappa value indicates a high level of agreement between the model's predictions and the ground truth labels, substantially surpassing the level expected by random chance. This demonstrates that the performance advantage of Infra-Net does not stem from biases in the training data distribution but is rather attributable to its stronger feature discrimination capability. Furthermore, it validates the robustness under class-imbalanced conditions.</p>
      <p id="d2e2931">In summary, experiments with different classification methods on this dataset demonstrate that Infra-Net consistently achieves the best performance, thereby validating its superior capability in multi-class interference infrasound signal recognition. Subsequently, the field-measured CTBTO dataset will be employed for evaluation.</p>
<sec id="Ch1.S4.SS3.SSS1">
  <label>4.3.1</label><title>Ablation experiment</title>
      <p id="d2e2941">To investigate the effect of the parallel structure, specifically the roles of the two branches and the confidence-based decision-making module in the classification of infrasound events, three ablation experiments are conducted: <list list-type="bullet"><list-item>
      <p id="d2e2946">Testing the classification capability of the GA-BiGRU branch itself: The entire MSCI-Net branch and the confidence-based decision-making module are removed, and only the GA-BiGRU branch is used for classification, denoted as Ma;</p></list-item><list-item>
      <p id="d2e2950">Testing the classification capability of the MSCI-Net branch itself: The entire GA-BiGRU branch and the confidence-based decision-making module are removed, and only the MSCI-Net branch is used for classification, denoted as Mb;</p></list-item><list-item>
      <p id="d2e2954">Testing the effectiveness of the dual-branch structure: Both the GA-BiGRU and MSCI-Net branches are retained, but the confidence-based decision-making module is replaced with a simple averaging operation, denoted as Mc;</p></list-item><list-item>
      <p id="d2e2958">Validating the effectiveness of the confidence-based decision-making module: The complete model proposed in the paper is used for classification, that is Infra-Net.</p></list-item></list> The radial basis function SVM (RBF-SVM) is established as the benchmark after demonstrating superior accuracy over alternative kernel functions in comparative tests. Corresponding ablation experiment results are presented in Table 7.</p>

<table-wrap id="T7"><label>Table 7</label><caption><p id="d2e2965">Results of ablation experiments.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Model</oasis:entry>
         <oasis:entry colname="col2">ACC (%)</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M80" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> (%)</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M81" display="inline"><mml:mrow><mml:mi>F</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Baseline</oasis:entry>
         <oasis:entry colname="col2">75.73</oasis:entry>
         <oasis:entry colname="col3">70.06</oasis:entry>
         <oasis:entry colname="col4">77.91</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Ma</oasis:entry>
         <oasis:entry colname="col2">79.59</oasis:entry>
         <oasis:entry colname="col3">75.95</oasis:entry>
         <oasis:entry colname="col4">80.44</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Mb</oasis:entry>
         <oasis:entry colname="col2">81.14</oasis:entry>
         <oasis:entry colname="col3">77.78</oasis:entry>
         <oasis:entry colname="col4">82.03</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Mc</oasis:entry>
         <oasis:entry colname="col2">81.89</oasis:entry>
         <oasis:entry colname="col3">78.05</oasis:entry>
         <oasis:entry colname="col4">82.73</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Infra-Net</oasis:entry>
         <oasis:entry colname="col2">82.07</oasis:entry>
         <oasis:entry colname="col3">78.62</oasis:entry>
         <oasis:entry colname="col4">82.76</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table><table-wrap-foot><p id="d2e2968">Note: ACC, the accuracy; <inline-formula><mml:math id="M79" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula>, precision.</p></table-wrap-foot></table-wrap>

      <p id="d2e3102">As can be observed from Table 7, all deep learning models demonstrate significantly superior performance compared to the SVM baseline (with the highest accuracy improvement reaching 6.34 %), proving the effectiveness of deep learning in the infrasound classification task. Furthermore, ablation experiment results indicate that the dual-branch parallel architecture (Mc) achieves better performance than any single-branch model (Ma or Mb), demonstrating that the feature extraction capabilities of the GA-BiGRU and MSCI-Net branches are complementary and that their fusion is necessary. Moreover, the confidence-based decision-making module proposed in this paper (Md) achieves optimal performance across all evaluation metrics, slightly outperforming the simple averaging fusion strategy (Mc). This indicates that the proposed module can more effectively integrate the predictive information from both branches, enabling more accurate decision-making. These results collectively validate the effectiveness of each module in the proposed architecture.</p>
      <p id="d2e3106">We also conducted a statistical significance test to more rigorously evaluate the performance difference between Infra-Net and the ablation models (especially Mc). The results are shown in Table 8.</p>

<table-wrap id="T8"><label>Table 8</label><caption><p id="d2e3112">The significance test results between Infra-Net and the closest-performing models such as Mc. The results show that the performance improvement of our method is statistically significant across multiple repeated experiments.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="5">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:colspec colnum="5" colname="col5" align="left"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Model</oasis:entry>
         <oasis:entry colname="col2">Source</oasis:entry>
         <oasis:entry colname="col3">Performance</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M82" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula>-value</oasis:entry>
         <oasis:entry colname="col5">Significant?</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Metric</oasis:entry>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Infra-Net vs. Mc</oasis:entry>
         <oasis:entry colname="col2">Table 7</oasis:entry>
         <oasis:entry colname="col3">ACC/<inline-formula><mml:math id="M83" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula>/<inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:mi>F</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M85" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M86" display="inline"><mml:mo>&lt;</mml:mo></mml:math></inline-formula> 0.05</oasis:entry>
         <oasis:entry colname="col5">Yes</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Infra-Net vs. Mb</oasis:entry>
         <oasis:entry colname="col2">Table 7</oasis:entry>
         <oasis:entry colname="col3">ACC/<inline-formula><mml:math id="M87" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula>/<inline-formula><mml:math id="M88" display="inline"><mml:mrow><mml:mi>F</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M89" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M90" display="inline"><mml:mo>&lt;</mml:mo></mml:math></inline-formula> 0.05</oasis:entry>
         <oasis:entry colname="col5">Yes</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Infra-Net vs. Ma</oasis:entry>
         <oasis:entry colname="col2">Table 7</oasis:entry>
         <oasis:entry colname="col3">ACC/<inline-formula><mml:math id="M91" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula>/<inline-formula><mml:math id="M92" display="inline"><mml:mrow><mml:mi>F</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M93" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M94" display="inline"><mml:mo>&lt;</mml:mo></mml:math></inline-formula> 0.05</oasis:entry>
         <oasis:entry colname="col5">Yes</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <fig id="F11" specific-use="star"><label>Figure 11</label><caption><p id="d2e3318">Confidence fusion process when Branch 1 is correct and Branch 2 is incorrect.</p></caption>
            <graphic xlink:href="https://nhess.copernicus.org/articles/26/4529/2026/nhess-26-4529-2026-f11.png"/>

          </fig>

      <fig id="F12" specific-use="star"><label>Figure 12</label><caption><p id="d2e3329">Confidence fusion process when Branch 1 is incorrect and Branch 2 is correct.</p></caption>
            <graphic xlink:href="https://nhess.copernicus.org/articles/26/4529/2026/nhess-26-4529-2026-f12.png"/>

          </fig>

      <p id="d2e3339">To intuitively demonstrate the effectiveness of this fusion mechanism, we have supplemented the visualization of the probability change process as you requested. It should be noted that when both branches make correct predictions or both make incorrect predictions, the fusion result is consistent with the branch results and cannot reflect the decision correcting capability of the fusion module. Therefore, we focus on representative samples where the two branches produce inconsistent classification conclusions. Figure 11 shows a sample where Branch 1 is correct and Branch 2 is incorrect; after inner product fusion, the joint confidence favors the true class, and the final output is a correct recognition result. Figure 12 shows the opposite case (Branch 1 incorrect, Branch 2 correct); the fusion module similarly corrects the influence of the erroneous branch and outputs a correct recognition result. These two cases fully demonstrate that the multi view learning strategy combined with the confidence-based fusion module proposed in this paper can effectively leverage the complementarity of the two branches in feature perspectives, suppress the adverse effects of single branch misjudgment, and thus significantly improve the accuracy and reliability of recognition results.</p>
</sec>
<sec id="Ch1.S4.SS3.SSS2">
  <label>4.3.2</label><title>Comparative experiment</title>
      <p id="d2e3350">Since the layer depth of the wavelet scattering network and the time-invariant scale can affect the dimensionality of the resulting wavelet scattering coefficients, which in turn influence the classification outcomes, comparative experiments were conducted to determine the optimal selection of network layers and time-invariant scales. These experiments aimed to derive general conclusions about the parameter settings. According to Wang et al. (2017), the optimal dimensionality of the feature matrix is achieved when the product of the scattering decomposition sampling frequency and the time-invariant scale is approximately equal to the total number of sampling points <inline-formula><mml:math id="M95" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>. Therefore, the time-invariant scale was primarily analyzed by considering cases in which it is equal to the signal duration and half the signal duration. Additionally, one-way analysis of variance (ANOVA-1) (Ai et al., 2008) was performed to verify that taking the logarithm of the wavelet scattering features can effectively enhance the separability. Finally, experiments were conducted by replacing the GA-BiGRU network in the Infra-Net with a GA-GRU network. This was done to validate the superior feature extraction capability of the GA-BiGRU compared to the GA-GRU. <list list-type="order"><list-item>
      <p id="d2e3362">Two-layer wavelet scattering networks were utilized, with the time-invariant scales set to 150 and 75 s. Infrasound signals with a data size of 1 <inline-formula><mml:math id="M96" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 3000 were input into the wavelet scattering networks to extract the logarithmic wavelet scattering features with data sizes of 6 <inline-formula><mml:math id="M97" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 168 and 12 <inline-formula><mml:math id="M98" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 123 as the inputs. These cases are referred to as methods Md1 and Md2.</p></list-item><list-item>
      <p id="d2e3387">A three-layer wavelet scattering network was employed, with the time-invariant scale set to 75 s (this study sets it to 150 s). Infrasound signals with a data size of 1 <inline-formula><mml:math id="M99" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 3000 were input into the wavelet scattering network to extract the logarithmic wavelet scattering features with a data size of 191 <inline-formula><mml:math id="M100" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 12 (this study yields a data size of 307 <inline-formula><mml:math id="M101" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 6) as inputs. This case is referred to as method Me.</p></list-item><list-item>
      <p id="d2e3412">The GA-BiGRU network was replaced with a GA-GRU network. This case is referred to as method Mf.</p></list-item></list> Figure 13 presents a comparison of the classification results for experiments (1)–(3).</p>

      <fig id="F13"><label>Figure 13</label><caption><p id="d2e3418">Comparison of the experimental results.</p></caption>
            <graphic xlink:href="https://nhess.copernicus.org/articles/26/4529/2026/nhess-26-4529-2026-f13.png"/>

          </fig>

      <p id="d2e3427">By comparing Md1 with Md2 and Me with the Infra-Net, it was found that when the number of network layers is fixed, the time-invariant scale directly affects the sample expansion multiple and the length of the data. Moreover, the data length of the scattering features has a direct impact on the classification results. Shorter data lengths are associated with poorer classification results, indicating that even though the number of extracted scattering feature paths grows exponentially, the amount of effective information they carry remains limited. Therefore, a balance must be struck between the dilation factor and the feature length. The experimental results lead to the final conclusion that setting the product of the scale parameter and the sampling frequency close to the total number of sampling points ensures sufficient information retention in the feature matrix while avoiding redundancy caused by excessive dimensionality, thus achieving better classification results. Additionally, it was found that when a three-layer network is used, the classification accuracy is superior to the case when a two-layer network is used, suggesting that a three-layer wavelet scattering network is generally more effective in extracting signal features. Finally, comparing Mf with the Infra-Net revealed that the GA-BiGRU network generally has an advantage over the GA-GRU network in terms of processing time series data.</p>
      <p id="d2e3431">In this study, after extracting the wavelet scattering features, their logarithm was taken to ascertain whether this transformation positively influenced the experimental outcomes. ANOVA-1 was conducted on the wavelet scattering features before and after taking the logarithm. The threshold for statistical significance was set to <inline-formula><mml:math id="M102" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M103" display="inline"><mml:mo>&lt;</mml:mo></mml:math></inline-formula> 0.05. This conventional level indicates that observed differences are unlikely to occur by chance.</p>

<table-wrap id="T9" specific-use="star"><label>Table 9</label><caption><p id="d2e3451">Results of the ANOVA-1 analysis before and after logarithmic transformation of the wavelet scattering features (bold font indicates enhanced significance of the differences).</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="7">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">Sample 1</oasis:entry>
         <oasis:entry colname="col3">Sample 2</oasis:entry>
         <oasis:entry colname="col4">Sample 3</oasis:entry>
         <oasis:entry colname="col5">Sample 4</oasis:entry>
         <oasis:entry colname="col6">Sample 5</oasis:entry>
         <oasis:entry colname="col7">Sample 6</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Before logarithm</oasis:entry>
         <oasis:entry colname="col2"><bold>0.0016</bold></oasis:entry>
         <oasis:entry colname="col3"><bold>0.0573</bold></oasis:entry>
         <oasis:entry colname="col4">0.0034</oasis:entry>
         <oasis:entry colname="col5"><bold>0.0182</bold></oasis:entry>
         <oasis:entry colname="col6"><bold>0.0807</bold></oasis:entry>
         <oasis:entry colname="col7"><bold>0.0518</bold></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">After logarithm</oasis:entry>
         <oasis:entry colname="col2"><bold>0.0009</bold></oasis:entry>
         <oasis:entry colname="col3"><bold>0.0016</bold></oasis:entry>
         <oasis:entry colname="col4">0.0950</oasis:entry>
         <oasis:entry colname="col5"><bold>0.0033</bold></oasis:entry>
         <oasis:entry colname="col6"><bold>6.42</bold> <inline-formula><mml:math id="M104" display="inline"><mml:mo mathvariant="bold">×</mml:mo></mml:math></inline-formula> <bold>10</bold><sup>−<bold>6</bold></sup></oasis:entry>
         <oasis:entry colname="col7"><bold>3.31</bold> <inline-formula><mml:math id="M106" display="inline"><mml:mo mathvariant="bold">×</mml:mo></mml:math></inline-formula> <bold>10</bold><sup>−<bold>9</bold></sup></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e3606">Features were extracted via a 150 s time-invariant scale scattering network. Each input raw signal (1 <inline-formula><mml:math id="M108" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 3000) generates six feature vectors (307 <inline-formula><mml:math id="M109" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> 1), denoted as Features 1 to 6. One-way ANOVA was performed on these six features both before and after logarithmic transformation, with results presented in Table 9. The comparisons reveal that with the exception of Feature 3, the significance of the differences was greater after the logarithm of the remaining Features was taken. Furthermore, there are three Features, which initially lacked significant differences, became significantly different after the logarithmic transformation. Additionally, random sampling analysis across all of the data confirmed the universality of this conclusion, thereby demonstrating that taking the logarithm of the wavelet scattering features, as in the proposed method, can enhance their separability and consequently improve the classification accuracy.</p>
      <p id="d2e3623">Furthermore, to substantiate the reliability of the classification performance of Infra-Net, comparative experiments have been conducted with more current advanced methods. These include advanced methods used in the field of infrasound signal classification, namely the CNN network (Bryan et al., 2018), the improved CNN-5 network (Tan et al., 2021), the improved LeNet-5 (Leng et al., 2022), the improved Alex-Net (Yuan et al., 2024), Prototype network (Zhao et al., 2024), and MS-SE-ResNet (Tan et al., 2024). The comparison also includes methods such as WST-AMResNet-18 (Liu et al., 2025), WST-Trans (Shirodkar et al., 2025), and WST-BiLSTM (Zhang et al., 2024b). For baseline comparison, SVM was employed as a traditional method.</p>

      <fig id="F14" specific-use="star"><label>Figure 14</label><caption><p id="d2e3628">Classification results for two types of infrasound events obtained using different models.</p></caption>
            <graphic xlink:href="https://nhess.copernicus.org/articles/26/4529/2026/nhess-26-4529-2026-f14.png"/>

          </fig>

      <p id="d2e3638">Figure 14 presents the classification results of the different models for two types of infrasound events. Compared with traditional methods, the classification accuracy of Infra-Net is improved by about 9 %. Even compared with the MS-SE-ResNet, which has the highest accuracy among them, Infra-Net still achieves an accuracy improvement of 3.55 %. When compared with other advanced methods related to the combination of wavelet scattering features and deep learning in this study, the accuracy is also significantly increased, which fully demonstrates that the confidence-based decision-making module has also played an important role in improving classification performance. It can also be observed that, compared with traditional infrasound classification methods, wavelet scattering based methods mostly achieve better classification results, indicating that it has great application potential in feature representation and small sample classification tasks, and will provide a reliable methodological and technical support for following research.</p>
</sec>
</sec>
</sec>
<sec id="Ch1.S5">
  <label>5</label><title>Discussion</title>
      <p id="d2e3652">This section discusses the performance differential between traditional wavelet scattering-based methods and the proposed approach, benchmarking it against several widely-adopted techniques.</p>
      <p id="d2e3655">The first comparative experiments were conducted: with fixed network architecture and parameters, event classification was performed using scattering features both before and after logarithmic processing as inputs, while simultaneously comparing the performance differences between column-wise decomposition of the scattering feature matrix into independent inputs versus traditional entire matrix input. Experimental results are detailed in Table 10.</p>

<table-wrap id="T10"><label>Table 10</label><caption><p id="d2e3661">Comparative Experimental Results.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="2">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Input method</oasis:entry>
         <oasis:entry colname="col2">ACC (%)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Infra-Net</oasis:entry>
         <oasis:entry colname="col2">82.07</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Without logarithmic processing</oasis:entry>
         <oasis:entry colname="col2">78.91</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Entire matrix input</oasis:entry>
         <oasis:entry colname="col2">77.89</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e3716">As shown in Table 10, the Infra-Net achieves superior accuracy compared to two conventional approaches. This result strongly validates that employing logarithmic transformation to amplify feature distinctions is an effective strategy for enhancing classification performance of low-frequency signals. Furthermore, the column-wise decomposition method for scattering matrices significantly outperforms traditional entire matrix input in accuracy. This advantage stems from a dual mechanism: the decomposition of scattering feature paths enhances training sufficiency, while information fusion theory enables comprehensive analysis of the results. Their synergistic effect significantly optimizes both classification accuracy and robustness.</p>
      <p id="d2e3719">To highlight the superiority of Infra-Net, in addition to the comparative experiments shown in Table 10, the classification accuracies of the LSTM, BiLSTM, WST-LSTM, WST-BiLSTM and WST-BiGRU models were also compared. Furthermore, to validate the efficacy of the purpose-built lightweight design in MSCI-Net for small-sample scenarios, the proposed structure was substituted with deeper network architectures, including Alex-Net (Krizhevsky et al., 2017), Visual Geometry Group (VGG)-16, and VGG-19, to compare the classification performances of the shallow and deep networks under small-sample conditions. The comparison results are presented in Table 11, and all results are the average values of each subset. The accuracy and other indices of the direct classification obtained using preprocessed data are significantly lower than those obtained using logarithmic wavelet scattering features, with improvements of nearly 20 % across all of the indices when these features are used, thus confirming their effectiveness. The results also indicate that replacing MSCI-Net with deeper networks did not enhance the accuracy. Although the classification efficiency values are not listed, the deep networks required significantly longer training time compared to MSCI-Net. This suggests that shallow networks can better and more rapidly learn generalized features from limited data in small-sample scenarios.</p>

<table-wrap id="T11"><label>Table 11</label><caption><p id="d2e3725">Classification results of the different methods.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Model</oasis:entry>
         <oasis:entry colname="col2">ACC</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M110" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M111" display="inline"><mml:mrow><mml:mi>F</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">(%)</oasis:entry>
         <oasis:entry colname="col3">(%)</oasis:entry>
         <oasis:entry colname="col4">(%)</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">LSTM</oasis:entry>
         <oasis:entry colname="col2">52.90</oasis:entry>
         <oasis:entry colname="col3">52.19</oasis:entry>
         <oasis:entry colname="col4">50.16</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">BiLSTM</oasis:entry>
         <oasis:entry colname="col2">55.11</oasis:entry>
         <oasis:entry colname="col3">54.97</oasis:entry>
         <oasis:entry colname="col4">50.89</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">WST-LSTM</oasis:entry>
         <oasis:entry colname="col2">77.38</oasis:entry>
         <oasis:entry colname="col3">75.42</oasis:entry>
         <oasis:entry colname="col4">77.49</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">WST-BiLSTM</oasis:entry>
         <oasis:entry colname="col2">77.33</oasis:entry>
         <oasis:entry colname="col3">76.24</oasis:entry>
         <oasis:entry colname="col4">77.41</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">WST-BiGRU</oasis:entry>
         <oasis:entry colname="col2">77.87</oasis:entry>
         <oasis:entry colname="col3">72.84</oasis:entry>
         <oasis:entry colname="col4">80.29</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Replacing MSCI-Net with Alex-Net</oasis:entry>
         <oasis:entry colname="col2">79.60</oasis:entry>
         <oasis:entry colname="col3">75.43</oasis:entry>
         <oasis:entry colname="col4">80.75</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Replacing MSCI-Net with VGG-16</oasis:entry>
         <oasis:entry colname="col2">81.19</oasis:entry>
         <oasis:entry colname="col3">76.96</oasis:entry>
         <oasis:entry colname="col4">82.16</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Replacing MSCI-Net with VGG-19</oasis:entry>
         <oasis:entry colname="col2">80.30</oasis:entry>
         <oasis:entry colname="col3">76.62</oasis:entry>
         <oasis:entry colname="col4">82.19</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Infra-Net</oasis:entry>
         <oasis:entry colname="col2">82.07</oasis:entry>
         <oasis:entry colname="col3">78.62</oasis:entry>
         <oasis:entry colname="col4">82.76</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e3926">As evidenced by the experimental results, the proposed method achieves a maximum classification accuracy of 82.07 % on the CTBTO measured dataset. Although this performance is lower than that obtained on the public LOTIS dataset, it maintains a clear advantage when compared horizontally with other approaches evaluated on the same CTBTO dataset. The performance discrepancy primarily stems from the CTBTO dataset's inherent characteristics: it contains infrasound events with a wide range of source distances and demonstrates greater susceptibility to noise interference in real monitoring environments, collectively increasing classification difficulty. Crucially, despite these challenges, the proposed method consistently delivers superior and stable classification performance across both distinct datasets, thereby validating its effectiveness and robustness in complex practical scenarios.</p>
</sec>
<sec id="Ch1.S6" sec-type="conclusions">
  <label>6</label><title>Conclusion</title>
      <p id="d2e3938">This study addresses the challenge of classifying infrasound signals with limited sample sizes in natural hazard monitoring through the Infra-Net model. Through systematic experimentation and analysis, key conclusions were drawn regarding data processing, model construction, and application efficacy, demonstrating its effectiveness in modeling complex systems for robust geophysical monitoring and disaster decision support.</p>
      <p id="d2e3941">In data processing, compared to conventional methods that use the entire matrix as input, this study improves classification accuracy by 8.3 %. Furthermore, using the extracted logarithmic wavelet scattering features as inputs, rather than raw signals, directly enhances classification precision by over 20 %. These results indicate that the data processing method significantly improves classification performance in small-sample scenarios.</p>
      <p id="d2e3944">In model construction, compared to the baseline model, Infra-Net improves all classification metrics by an average of 6.58 %, fully validating the effectiveness of its multi-branch collaborative modeling strategy in enhancing the discriminative capability of the model.</p>
      <p id="d2e3947">In application efficacy, based on the conclusion that the optimal feature dimension satisfies the parameter setting of scattering sampling frequency <inline-formula><mml:math id="M112" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> time-invariant scale <inline-formula><mml:math id="M113" display="inline"><mml:mo>≈</mml:mo></mml:math></inline-formula> signal sampling points, this method is expected to achieve predictable and robust performance in analogous application scenarios. Despite the above results, this study still has several limitations that urgently need to be addressed in future work.</p>
      <p id="d2e3965">First, the performance gap between the public LOTIS dataset and the CTBTO field-measured data indicates that the current model's recognition capability is significantly constrained when dealing with signal distortions caused by extremely long-range propagation, and its cross-scene generalization ability still needs improvement. Furthermore, the ablation experiment results show that the confidence-based fusion strategy proposed in this paper achieves only limited improvement over simple averaging fusion. To address these limitations, we will focus on the following research directions. First of all, we will introduce domain adaptation strategies, for example, reducing the model's dependence on specific monitoring environments through transfer learning to narrow the generalization gap. Second, we will collect and annotate more diverse infrasound event data (including various natural and anthropogenic events at different propagation distances and signal-to-noise ratios) to increase the coverage of training samples. Third, we will explore uncertainty-aware fusion methods to better handle conflicts among predictions from different views, thereby improving decision robustness and supporting more reliable natural hazard early warning.</p>
</sec>

      
      </body>
    <back><notes notes-type="dataavailability"><title>Data availability</title>

      <p id="d2e3973">The datasets supporting the findings of this study are available from the corresponding author upon reasonable request. The LOTIS dataset can be accessed through the associated publication at <ext-link xlink:href="https://doi.org/10.3390/acoustics8020021" ext-link-type="DOI">10.3390/acoustics8020021</ext-link> (Yin et al., 2026). The CTBTO infrasound data are available via the official website (<uri>https://www.ctbto.org/specials/vdec/</uri>, last access: 21 September 2026) of the Preparatory Commission for the CTBTO subject to their standard application and authorization procedures.</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d2e3985">Hongru Li: Conceptualization, Methodology, Software, Writing-Original draft preparation; Xihai Li: Data curation; Jihao Liu: Supervision; Shengjie Luo: Software; Yun Zhang: Software, Validation.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d2e3991">The contact author has declared that none of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d2e3997">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. The authors bear the ultimate responsibility for providing appropriate place names. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.</p>
  </notes><ack><title>Acknowledgements</title><p id="d2e4003">I would like to express my sincere gratitude to Professor Yang Jun and his team for their significant assistance with the data aspect of this study.</p></ack><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d2e4008">This paper was edited by Maneesha Vinodini Ramesh and reviewed by two anonymous referees.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bib1"><label>1</label><mixed-citation>Abbaszadeh, M., Soltani, S. M., and Ahmed, A. N.: Optimization of support vector machine parameters in modeling of Iju deposit mineralization and alteration zones using particle swarm optimization algorithm and grid search method, Comput. Geosci., 165, 105140, <ext-link xlink:href="https://doi.org/10.1016/j.cageo.2022.105140" ext-link-type="DOI">10.1016/j.cageo.2022.105140</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bib2"><label>2</label><mixed-citation>Ai, L., Wang, J., and Wang, X.: Multi-features fusion diagnosis of tremor based on artificial neural network and D–S evidence theory, Signal. Process., 88, 2927–2935, <ext-link xlink:href="https://doi.org/10.1016/j.sigpro.2008.06.018" ext-link-type="DOI">10.1016/j.sigpro.2008.06.018</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bib3"><label>3</label><mixed-citation>Al, J. A. and Khushaba, R. N.: Deep hand gesture recognition: a wavelet scattering alternative to convolutional networks, in: 2023 IEEE Statistical Signal Processing Workshop (SSP), Hanoi, Vietnam, 2–5 July 2023, 438–442, <ext-link xlink:href="https://doi.org/10.1109/SSP53291.2023.10208011" ext-link-type="DOI">10.1109/SSP53291.2023.10208011</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bib4"><label>4</label><mixed-citation>Albert, S. and Linville, L.: Benchmarking current and emerging approaches to infrasound signal classification, Seismol. Res. Lett., 91, 921–929, <ext-link xlink:href="https://doi.org/10.1785/0220190116" ext-link-type="DOI">10.1785/0220190116</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib5"><label>5</label><mixed-citation>Alegria, O. C., Valtierra-Rodriguez, M., P. Amezquita-Sanchez, J., Millan-Almaraz, J. R., Rodriguez, L. M., Moctezuma, A. M., Dominguez-Gonzalez, A., and Cruz-Abeyro, J. A.: Empirical wavelet transform-based detection of anomalies in ULF geomagnetic signals associated to seismic events with a fuzzy logic-based system for automatic diagnosis, in: Wavelet Transform and Some of Its Real-World Applications, edited by: Baleanu, D., InTech, 111–124, <ext-link xlink:href="https://doi.org/10.5772/61163" ext-link-type="DOI">10.5772/61163</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bib6"><label>6</label><mixed-citation>Andén, J. and Mallat, S.: Deep scattering spectrum, IEEE Trans. Signal. Process., 62, 4114–2418, <ext-link xlink:href="https://doi.org/10.1109/TSP.2014.2326991" ext-link-type="DOI">10.1109/TSP.2014.2326991</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bib7"><label>7</label><mixed-citation>Bruna, J. and Mallat, S.: Classification with scattering operators, in: CVPR 2011, Colorado Springs, CO, USA, 20–25 June 2011, 1561–1566, <ext-link xlink:href="https://doi.org/10.1109/CVPR.2011.5995635" ext-link-type="DOI">10.1109/CVPR.2011.5995635</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bib8"><label>8</label><mixed-citation>Bruna, J. and Mallat, S.: Invariant scattering convolution networks, IEEE Trans. Pattern. Anal. Mach. Intell., 35, 1872–1886, <ext-link xlink:href="https://doi.org/10.1109/tpami.2012.230" ext-link-type="DOI">10.1109/tpami.2012.230</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bib9"><label>9</label><mixed-citation>Bryan, K. J., Smith, K. E., Solomon, M. L.,  Clauter, D. A.,  Smith, A. O., and  Peter, A. M.: Deep wavelet scattering features for infrasonic threat identification, in: Proceedings of the Chemical, Biological, Radiological, Nuclear, and Explosives Sensing XIX, Orlando, FL, United States, 16 May 2018, 1062901–1062918, <ext-link xlink:href="https://doi.org/10.1117/12.2304544" ext-link-type="DOI">10.1117/12.2304544</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib10"><label>10</label><mixed-citation>Buriro, A. B., Ahmed, B., Baloch, G., Ahmed, J., Shoorangiz, R., Weddell, S. J., and Jones, R. D.: Classification of alcoholic eeg signals using wavelet scattering transform-based features, Comput. Biol. Med., 139, 104969, <ext-link xlink:href="https://doi.org/10.1016/j.compbiomed.2021.104969" ext-link-type="DOI">10.1016/j.compbiomed.2021.104969</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib11"><label>11</label><mixed-citation>Chang, T. Y. and Cui, H. L.: Determination of direction of arrival of seismic wave by a single Tri-axial fiber optic geophone, Def. Technol., 9, 1–9, <ext-link xlink:href="https://doi.org/10.1016/j.dt.2013.02.001" ext-link-type="DOI">10.1016/j.dt.2013.02.001</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bib12"><label>12</label><mixed-citation>Cho, K., Merrienboer, B. V., Gulcehre, C., Bougares, F., Schwenk, H., and Bengio, Y.: Learning phrase representations using RNN encoder–decoder for statistical machine translation, in: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar, 3 June 2014, 1724–1734, <ext-link xlink:href="https://doi.org/10.3115/v1/D14-1179" ext-link-type="DOI">10.3115/v1/D14-1179</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bib13"><label>13</label><mixed-citation>Dai, R., Wang, Z., Wang, W. L., Jie, L., Chen, J. C., and Ye, Q. L.: VTNet: A multi-domain information fusion model for long-term multi-variate time series forecasting with application in irrigation water level, Appl. Soft. Comput., 167, 112251, <ext-link xlink:href="https://doi.org/10.1016/j.asoc.2024.112251" ext-link-type="DOI">10.1016/j.asoc.2024.112251</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bib14"><label>14</label><mixed-citation> Dai, Y. J., Teng, P. X., Lyu, J., Ji, P. F., and Cheng, W.: Analysis of the infrasound signals from rocket launch, J. Appl. Acoust., 40, 676–683, 2021.</mixed-citation></ref>
      <ref id="bib1.bib15"><label>15</label><mixed-citation>Ding, Z. W., Zhang, C. F., Huang, X., Liu, Q. S., Liu, B., Gao, F., Li, L., and Liu, Y. X.: Recognition method of coal–rock reflection spectrum using wavelet scattering transform and bidirectional Long-Short-Term Memory, Rock. Mech. Rock. Eng., 57, 1353–1374, <ext-link xlink:href="https://doi.org/10.1007/s00603-023-03600-z" ext-link-type="DOI">10.1007/s00603-023-03600-z</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bib16"><label>16</label><mixed-citation> Fan, X., Cheng, J. Y., Wang, Y. H., Li, S., Duan, J. H., and Wang, P.: Intelligent recognition of coal mine microseismic signal based on wavelet scattering decomposition transform, J. China. Coal. Soc., 47, 2722–2731, 2022.</mixed-citation></ref>
      <ref id="bib1.bib17"><label>17</label><mixed-citation>Hochreiter, S. and Schmidhuber, J.: Long short-term memory, Neural. Comput., 9, 1735–1780, <ext-link xlink:href="https://doi.org/10.1162/neco.1997.9.8.1735" ext-link-type="DOI">10.1162/neco.1997.9.8.1735</ext-link>, 1997.</mixed-citation></ref>
      <ref id="bib1.bib18"><label>18</label><mixed-citation>Huang, W., Zou, X. Y., and Qiao, R.: Scattering wavelet based deep network for image classification, in: Communications in Computer and Information Science, edited by: Li, G., Filipe, J., and Xu, Z. W., Springer Singapore, 459–469, <ext-link xlink:href="https://doi.org/10.1007/978-981-10-8530-7_45" ext-link-type="DOI">10.1007/978-981-10-8530-7_45</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib19"><label>19</label><mixed-citation>Khan, A. A., Dhawan, A., Akhlaghi, N., Majdi, J. A., and Sikdar, S.: Application of wavelet scattering networks in classification of ultrasound image sequences, in: Proceedings of the 2017 IEEE International Ultrasonics Symposium (IUS), Washington, DC, USA, 6–9 September 2017, 1–4, <ext-link xlink:href="https://doi.org/10.1109/ULTSYM.2017.8091649" ext-link-type="DOI">10.1109/ULTSYM.2017.8091649</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib20"><label>20</label><mixed-citation>Koch, K. and Pilger, C.: Infrasound observations from the site of past underground nuclear explosions in North Korea, Geophys. J. Int., 216, 182–200, <ext-link xlink:href="https://doi.org/10.1093/gji/ggy381" ext-link-type="DOI">10.1093/gji/ggy381</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib21"><label>21</label><mixed-citation>Kogelnig, A., Hübl, J., Suriñach, E., Vilajosana, I., and McArdell, B. W.: Infrasound produced by debris flow: propagation and frequency content evolution, Nat. Hazard., 70, 1713–1733, <ext-link xlink:href="https://doi.org/10.1007/s11069-011-9741-8" ext-link-type="DOI">10.1007/s11069-011-9741-8</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bib22"><label>22</label><mixed-citation>Krizhevsky, A., Sutskever, I., and Hinton, G. E.: ImageNet classification with deep convolutional neural networks, J. Commun. ACM, 60, 84–90, <ext-link xlink:href="https://doi.org/10.1145/3065386" ext-link-type="DOI">10.1145/3065386</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib23"><label>23</label><mixed-citation>Lee, S. and Hong, S. G.: Classification of nuclear activity types for neighboring countries of South Korea using machine learning techniques with xenon isotopic activity ratios, Nucl. Eng. Technol., 56, 1372–1384, <ext-link xlink:href="https://doi.org/10.1016/j.net.2023.11.042" ext-link-type="DOI">10.1016/j.net.2023.11.042</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bib24"><label>24</label><mixed-citation>Leng X. P., Feng, L. Y., Ou, O., Du, X. L., Liu, D. L., and Tang, X.: Debris flow infrasound recognition method based on improved LeNet-5 network, Sustainability-Basel, 14, 15925, <ext-link xlink:href="https://doi.org/10.3390/su142315925" ext-link-type="DOI">10.3390/su142315925</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bib25"><label>25</label><mixed-citation>Li, H. R., Li, X. H., Tan, X. F., Liu, T. Y., Zhang, Y., Niu, C., and Liu, J. H.: Infrasound event classification fusion model based on multiscale SE-CNN and BiLSTM, Appl. Geophys., 21, 579–592, <ext-link xlink:href="https://doi.org/10.1007/s11770-024-1089-4" ext-link-type="DOI">10.1007/s11770-024-1089-4</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bib26"><label>26</label><mixed-citation>Li, H. R., Li, X. H., Liu, J. H., Wang, Y. T., Liu, Z. G., and Zeng, X. N.: Machine learning-driven classification of natural disasters via parallel confidence fusion, Mach. Learn.: Sci. Technol., 7, 015028, <ext-link xlink:href="https://doi.org/10.1088/2632-2153/ae3c58" ext-link-type="DOI">10.1088/2632-2153/ae3c58</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bib27"><label>27</label><mixed-citation>Li, J. H., Ke, L., Du, Q., Ding, X. D., Chen, X. M., and Wang, D. N.: Heart sound signal classification algorithm: a combination of wavelet scattering transform and twin Support Vector Machine, IEEE Access, 7, 179339–179348, <ext-link xlink:href="https://doi.org/10.1109/access.2019.2959081" ext-link-type="DOI">10.1109/access.2019.2959081</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib28"><label>28</label><mixed-citation>Li, M., Liu, X., and Liu, X.: Infrasound signal classification based on spectral entropy and support vector machine, Appl. Acoust., 113, 116–120, <ext-link xlink:href="https://doi.org/10.1016/j.apacoust.2016.06.019" ext-link-type="DOI">10.1016/j.apacoust.2016.06.019</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bib29"><label>29</label><mixed-citation>Lilly, J. M. and Olhede, S. C.: On the analytic wavelet transform, IEEE Trans. Inf. Theory., 56, 4135–4156, <ext-link xlink:href="https://doi.org/10.1109/tit.2010.2050935" ext-link-type="DOI">10.1109/tit.2010.2050935</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bib30"><label>30</label><mixed-citation>Liu, L., Wu, J., Li, D., Senhadji, L., and Shu, H.: Fractional wavelet scattering network and applications, IEEE Trans. Biomed. Eng., 66, 553–563, <ext-link xlink:href="https://doi.org/10.1109/tbme.2018.2850356" ext-link-type="DOI">10.1109/tbme.2018.2850356</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib31"><label>31</label><mixed-citation>Liu, X., Li, M., Tang, W., Wang, S., and Wu, X.: A new classification method of infrasound events using Hilbert-Huang Transform and Support Vector Machine, Math. Probl. Eng., 2014, 1–6, <ext-link xlink:href="https://doi.org/10.1155/2014/456818" ext-link-type="DOI">10.1155/2014/456818</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bib32"><label>32</label><mixed-citation>Liu, Y. X., Zhang, B. Q., Kong, F. T., Wang, B., Luo, C. M., and Ma, L.: Underwater acoustic classification using wavelet scattering transform and convolutional neural network with limited dataset, Appl. Acoust., 232, 110564, <ext-link xlink:href="https://doi.org/10.1016/j.apacoust.2025.110564" ext-link-type="DOI">10.1016/j.apacoust.2025.110564</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bib33"><label>33</label><mixed-citation>Lone, A. W. and Aydin, N.: Wavelet scattering transform based doppler signal classification, Comput. Biol. Med., 167, 107611, <ext-link xlink:href="https://doi.org/10.1016/j.compbiomed.2023.107611" ext-link-type="DOI">10.1016/j.compbiomed.2023.107611</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bib34"><label>34</label><mixed-citation>Lu, X., Liu, S., Chen, X., and Zhang, K.: Identification of the lubrication state of journal bearings based on acoustic emission and WST-CNN collaboration, J. Vib. Shock., 42, 71–77, <ext-link xlink:href="https://doi.org/10.13465/j.cnki.jvs.2023.22.008" ext-link-type="DOI">10.13465/j.cnki.jvs.2023.22.008</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bib35"><label>35</label><mixed-citation>Luong, T., Pham, H., and Manning, C. D.: Effective approaches to attention-based neural machine translation, in: Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP), Lisbon, Portugal, 17 August 2015, 1412–1421, <ext-link xlink:href="https://doi.org/10.18653/v1/D15-1166" ext-link-type="DOI">10.18653/v1/D15-1166</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bib36"><label>36</label><mixed-citation>Marchetti, E., Walter, F., Barfucci, G., Genco, R., Wenner, M., Ripepe, M., McArdell, B., and Price, C.: Infrasound array analysis of debris flow activity and implication for early warning, J. Geophys. Res.-Earth, 124, 567–587, <ext-link xlink:href="https://doi.org/10.1029/2018jf004785" ext-link-type="DOI">10.1029/2018jf004785</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib37"><label>37</label><mixed-citation>Park, J., Che, I. Y., Stump, B., Hayward, C., Dannemann, F., SeongJu, J., Kwong, K., McComas, S., Oldham, H. R., Scales, M. M., and Wright, V.: Characteristics of infrasound signals from North Korean underground nuclear explosions on 2016 January 6 and September 9, Geophys. J. Int., 214, 1865–1885, <ext-link xlink:href="https://doi.org/10.1093/gji/ggy252" ext-link-type="DOI">10.1093/gji/ggy252</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib38"><label>38</label><mixed-citation>Pásztor, M., Czanik, C., and Bondár, I.: A single array approach for infrasound signal discrimination from quarry blasts via machine learning, Remote. Sens., 15, 1657, <ext-link xlink:href="https://doi.org/10.3390/rs15061657" ext-link-type="DOI">10.3390/rs15061657</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bib39"><label>39</label><mixed-citation>Priya, B. L., Jayalakshmy, S., Pragatheeswaran, J. K., Saraswathi, D., and Poonguzhali, N.: Scattering convolutional network based predictive model for cognitive activity of brain using empirical wavelet decomposition, Biomed. Signal. Process. Control., 66, 102501, <ext-link xlink:href="https://doi.org/10.1016/j.bspc.2021.102501" ext-link-type="DOI">10.1016/j.bspc.2021.102501</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib40"><label>40</label><mixed-citation>Shirodkar, V. R., Edla, D. R., and Kumari, A.: Advancing motor imagery EEG classification through wavelet scattering transforms and 1D transformers, Procedia Comput. Sci., 258, 2860–2869, <ext-link xlink:href="https://doi.org/10.1016/j.procs.2025.04.546" ext-link-type="DOI">10.1016/j.procs.2025.04.546</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bib41"><label>41</label><mixed-citation>Singhal, S. and Kumar, M.: SPTDMD-WST: Arrhythmia classification from spatiotemporal modes of dynamic mode decomposition using wavelet scattering transform, Biomed. Signal. Process. Control., 92, 105983, <ext-link xlink:href="https://doi.org/10.1016/j.bspc.2024.105983" ext-link-type="DOI">10.1016/j.bspc.2024.105983</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bib42"><label>42</label><mixed-citation>Stéphane, M.: CHAPTER 7-Wavelet Bases, in: A Wavelet Tour of Signal Processing, Academic Press, 263–376, <ext-link xlink:href="https://doi.org/10.1016/B978-0-12-374370-1.00011-2" ext-link-type="DOI">10.1016/B978-0-12-374370-1.00011-2</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bib43"><label>43</label><mixed-citation>Souli, S., Amami, R., and Yahia, S. B.: A robust pathological voices recognition system based on DCNN and scattering transform, Appl. Acoust., 177, 107854, <ext-link xlink:href="https://doi.org/10.1016/j.apacoust.2020.107854" ext-link-type="DOI">10.1016/j.apacoust.2020.107854</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib44"><label>44</label><mixed-citation>Tan, X. F., Li, X. H., Liu, J. H., Li, G., and Yu, X. T.: Classification of chemical explosion and earthquake infrasound based on 1-D convolutional neural network, J. Appl. Acoust., 40, 457–467, <ext-link xlink:href="https://doi.org/10.11684/j.issn.1000-310X.2021.03.018" ext-link-type="DOI">10.11684/j.issn.1000-310X.2021.03.018</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib45"><label>45</label><mixed-citation>Tan, X. F., Li, X. H., Niu, C., Zeng, X. N., Li, H, R., and Liu, T. Y.: Classification method of infrasound events based on the MVIDA algorithm and MS-SE-ResNet, Appl. Geophys., 21, 667–679, <ext-link xlink:href="https://doi.org/10.1007/s11770-024-1112-9" ext-link-type="DOI">10.1007/s11770-024-1112-9</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bib46"><label>46</label><mixed-citation>Thüring, T., Schoch, M., Herwijnen, A., and Schweizer, J.: Robust snow avalanche detection using supervised machine learning with infrasonic sensor arrays, Cold. Reg. Sci. Technol., 111, 60–66, <ext-link xlink:href="https://doi.org/10.1016/j.coldregions.2014.12.014" ext-link-type="DOI">10.1016/j.coldregions.2014.12.014</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bib47"><label>47</label><mixed-citation>Toney, L., Fee, D., Witsil, A., and Matoza, R. S.: Waveform features strongly control subcrater classification performance for a large labeled volcano infrasound dataset, Seism. Rec., 2, 167–175, <ext-link xlink:href="https://doi.org/10.1785/0320220019" ext-link-type="DOI">10.1785/0320220019</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bib48"><label>48</label><mixed-citation>Tuan, D. P.: Classification of motor-imagery tasks using a large eeg dataset by fusing classifiers learning on wavelet-scattering features, IEEE Trans. Neural. Syst. Rehabil. Eng., 31, 1097–1107, <ext-link xlink:href="https://doi.org/10.1109/TNSRE.2023.3241241" ext-link-type="DOI">10.1109/TNSRE.2023.3241241</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bib49"><label>49</label><mixed-citation>Wang, H., Li, S., Zhou, Y., and Chen, S.: SAR automatic target recognition using a roto-translational invariant wavelet-scattering convolution network, Remote Sens., 10, 501, <ext-link xlink:href="https://doi.org/10.3390/rs10040501" ext-link-type="DOI">10.3390/rs10040501</ext-link>, 2018a.</mixed-citation></ref>
      <ref id="bib1.bib50"><label>50</label><mixed-citation>Wang, L., Liu, B., Wang, H., Zhong, G., and Dong, J.: Deep gabor scattering network for image classification, in: Pattern Recognition and Computer Vision, 2 November 2018, Springer International Publishing, 332–343, <ext-link xlink:href="https://doi.org/10.1007/978-3-030-03335-4_29" ext-link-type="DOI">10.1007/978-3-030-03335-4_29</ext-link>, 2018b.</mixed-citation></ref>
      <ref id="bib1.bib51"><label>51</label><mixed-citation>Wang, W., Gao, X., and Le, L.: Study of the similarities in scale models of a single-layer spherical lattice shell structure under the effect of internal explosion, Shock. Vib., 4, 1–13, <ext-link xlink:href="https://doi.org/10.1155/2017/9181729" ext-link-type="DOI">10.1155/2017/9181729</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib52"><label>52</label><mixed-citation>Wang, Z. X., Gu, W. B., Liang, T., Zhao, S. T., Chen, P., and Yu, L. F.: Monitoring and prediction of the vibration intensity of seismic waves induced in underwater rock by underwater drilling and blasting, Def. Technol., 18, 109–118, <ext-link xlink:href="https://doi.org/10.1016/j.dt.2020.10.007" ext-link-type="DOI">10.1016/j.dt.2020.10.007</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bib53"><label>53</label><mixed-citation>Wen, Y. D., Ren, W. T., Yang, Y., Chen, J., Han, Y., and Ye, Y. Z.: Analysis of infrasound signals generated by CZ-7 teleport 2 launch vehicle, J. Ordnance Equip. Eng., 40, 70–73, <ext-link xlink:href="https://doi.org/10.11809/bqzbgcxb2019.09.015" ext-link-type="DOI">10.11809/bqzbgcxb2019.09.015</ext-link>, 2019. </mixed-citation></ref>
      <ref id="bib1.bib54"><label>54</label><mixed-citation>Wiatowski, T. and Bolcskei, H.: A mathematical theory of deep convolutional neural networks for feature extraction, IEEE Trans. Inf. Theory., 64, 1845–1866, <ext-link xlink:href="https://doi.org/10.1109/TIT.2017.2776228" ext-link-type="DOI">10.1109/TIT.2017.2776228</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib55"><label>55</label><mixed-citation>Yin, H., Lu, Y., Wu, Y. H., Cheng, W., Pang, X. L., and Li, P.: Infrasound signal classification fusion model based on double-branch and multi-scale CNN and LSTM, Acoustics, 8, 21, <ext-link xlink:href="https://doi.org/10.3390/acoustics8020021" ext-link-type="DOI">10.3390/acoustics8020021</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bib56"><label>56</label><mixed-citation>Yuan, L., Liu, D. L., Sang, X. J., Zhang, S. J., and Chen, Q.: Debris flow infrasound signal recognition approach based on improved AlexNet, Computer and Modernization, 3, 1–6, <ext-link xlink:href="https://doi.org/10.3969/j.issn.1006-2475.2024.03.001" ext-link-type="DOI">10.3969/j.issn.1006-2475.2024.03.001</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bib57"><label>57</label><mixed-citation>Zhang, A., Chun, S., Cheng, Z., and Zhao, P.: Predicting the core thermal hydraulic parameters with a gated recurrent unit model based on the soft attention mechanism, Nucl. Eng. Technol., 56, 2343–2351, <ext-link xlink:href="https://doi.org/10.1016/j.net.2024.01.045" ext-link-type="DOI">10.1016/j.net.2024.01.045</ext-link>, 2024a.</mixed-citation></ref>
      <ref id="bib1.bib58"><label>58</label><mixed-citation>Zhang, H. Y., Zhao, Z. J., Liu, C., Duan, M., Lu, Z. G., and Wang, H.: Classification of motor imagery EEG signals using wavelet scattering transform and Bi-directional long short-term memory networks, Biocybern. Biomed. Eng., 44, 874–884, <ext-link xlink:href="https://doi.org/10.1016/j.bbe.2024.11.003" ext-link-type="DOI">10.1016/j.bbe.2024.11.003</ext-link>, 2024b.</mixed-citation></ref>
      <ref id="bib1.bib59"><label>59</label><mixed-citation>Zhang, Q., Wang, P., Pedrycz, W., and Li, Z.: Neighborhood entropy guided by a decision attribute and its applications in multi-source information fusion and attribute selection, Appl. Soft. Comput., 167, 112380, <ext-link xlink:href="https://doi.org/10.1016/j.asoc.2024.112380" ext-link-type="DOI">10.1016/j.asoc.2024.112380</ext-link>, 2024c.</mixed-citation></ref>
      <ref id="bib1.bib60"><label>60</label><mixed-citation>Zhang, X., Shen, J., Li, J., and Liu, X.: An instance-oriented multi-source information fusion technique based on neighborhood granules, Appl. Soft. Comput., 181, 113483, <ext-link xlink:href="https://doi.org/10.1016/j.asoc.2025.113483" ext-link-type="DOI">10.1016/j.asoc.2025.113483</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bib61"><label>61</label><mixed-citation> Zhao, Z. J., Chen, W., Ji, P. F., Teng, P. X., Lyu, J., and Yang, J.: A method for classification of few-shot infrasound signals applying prototype network, J. Appl. Acoust., 43, 1193–1202, 2024.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Infra-Net: a robust parallel decision-making network for discriminating natural hazards and anthropogenic infrasound events via multi-view feature learning</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>1</label><mixed-citation>
      
Abbaszadeh, M., Soltani, S. M., and Ahmed, A. N.: Optimization of support vector machine parameters in modeling of Iju deposit mineralization and alteration zones using particle swarm optimization algorithm and grid search method, Comput. Geosci., 165, 105140, <a href="https://doi.org/10.1016/j.cageo.2022.105140" target="_blank">https://doi.org/10.1016/j.cageo.2022.105140</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>2</label><mixed-citation>
      
Ai, L., Wang, J., and Wang, X.: Multi-features fusion diagnosis of tremor based on artificial neural network and D–S evidence theory, Signal. Process., 88, 2927–2935, <a href="https://doi.org/10.1016/j.sigpro.2008.06.018" target="_blank">https://doi.org/10.1016/j.sigpro.2008.06.018</a>, 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>3</label><mixed-citation>
      
Al, J. A. and Khushaba, R. N.: Deep hand gesture recognition: a wavelet scattering alternative to convolutional networks, in: 2023 IEEE Statistical Signal Processing Workshop (SSP), Hanoi, Vietnam, 2–5 July 2023, 438–442, <a href="https://doi.org/10.1109/SSP53291.2023.10208011" target="_blank">https://doi.org/10.1109/SSP53291.2023.10208011</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>4</label><mixed-citation>
      
Albert, S. and Linville, L.: Benchmarking current and emerging approaches to infrasound signal classification, Seismol. Res. Lett., 91, 921–929, <a href="https://doi.org/10.1785/0220190116" target="_blank">https://doi.org/10.1785/0220190116</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>5</label><mixed-citation>
      
Alegria, O. C., Valtierra-Rodriguez, M., P. Amezquita-Sanchez, J., Millan-Almaraz, J. R., Rodriguez, L. M., Moctezuma, A. M., Dominguez-Gonzalez, A., and Cruz-Abeyro, J. A.: Empirical wavelet transform-based detection of anomalies in ULF geomagnetic signals associated to seismic events with a fuzzy logic-based system for automatic diagnosis, in: Wavelet Transform and Some of Its Real-World Applications, edited by: Baleanu, D., InTech, 111–124, <a href="https://doi.org/10.5772/61163" target="_blank">https://doi.org/10.5772/61163</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>6</label><mixed-citation>
      
Andén, J. and Mallat, S.: Deep scattering spectrum, IEEE Trans. Signal. Process., 62, 4114–2418, <a href="https://doi.org/10.1109/TSP.2014.2326991" target="_blank">https://doi.org/10.1109/TSP.2014.2326991</a>, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>7</label><mixed-citation>
      
Bruna, J. and Mallat, S.: Classification with scattering operators, in: CVPR 2011, Colorado Springs, CO, USA, 20–25 June 2011, 1561–1566, <a href="https://doi.org/10.1109/CVPR.2011.5995635" target="_blank">https://doi.org/10.1109/CVPR.2011.5995635</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>8</label><mixed-citation>
      
Bruna, J. and Mallat, S.: Invariant scattering convolution networks, IEEE Trans. Pattern. Anal. Mach. Intell., 35, 1872–1886, <a href="https://doi.org/10.1109/tpami.2012.230" target="_blank">https://doi.org/10.1109/tpami.2012.230</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>9</label><mixed-citation>
      
Bryan, K. J., Smith, K. E., Solomon, M. L.,  Clauter, D. A.,  Smith, A. O., and  Peter, A. M.: Deep wavelet scattering features for infrasonic threat identification, in: Proceedings of the Chemical, Biological, Radiological, Nuclear, and Explosives Sensing XIX, Orlando, FL, United States, 16 May 2018, 1062901–1062918, <a href="https://doi.org/10.1117/12.2304544" target="_blank">https://doi.org/10.1117/12.2304544</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>10</label><mixed-citation>
      
Buriro, A. B., Ahmed, B., Baloch, G., Ahmed, J., Shoorangiz, R., Weddell, S. J., and Jones, R. D.: Classification of alcoholic eeg signals using wavelet scattering transform-based features, Comput. Biol. Med., 139, 104969, <a href="https://doi.org/10.1016/j.compbiomed.2021.104969" target="_blank">https://doi.org/10.1016/j.compbiomed.2021.104969</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>11</label><mixed-citation>
      
Chang, T. Y. and Cui, H. L.: Determination of direction of arrival of seismic wave by a single Tri-axial fiber optic geophone, Def. Technol., 9, 1–9, <a href="https://doi.org/10.1016/j.dt.2013.02.001" target="_blank">https://doi.org/10.1016/j.dt.2013.02.001</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>12</label><mixed-citation>
      
Cho, K., Merrienboer, B. V., Gulcehre, C., Bougares, F., Schwenk, H., and Bengio, Y.: Learning phrase representations using RNN encoder–decoder for statistical machine translation, in: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar, 3 June 2014, 1724–1734, <a href="https://doi.org/10.3115/v1/D14-1179" target="_blank">https://doi.org/10.3115/v1/D14-1179</a>, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>13</label><mixed-citation>
      
Dai, R., Wang, Z., Wang, W. L., Jie, L., Chen, J. C., and Ye, Q. L.: VTNet: A multi-domain information fusion model for long-term multi-variate time series forecasting with application in irrigation water level, Appl. Soft. Comput., 167, 112251, <a href="https://doi.org/10.1016/j.asoc.2024.112251" target="_blank">https://doi.org/10.1016/j.asoc.2024.112251</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>14</label><mixed-citation>
      
Dai, Y. J., Teng, P. X., Lyu, J., Ji, P. F., and Cheng, W.: Analysis of the infrasound signals from rocket launch, J. Appl. Acoust., 40, 676–683, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>15</label><mixed-citation>
      
Ding, Z. W., Zhang, C. F., Huang, X., Liu, Q. S., Liu, B., Gao, F., Li, L., and Liu, Y. X.: Recognition method of coal–rock reflection spectrum using wavelet scattering transform and bidirectional Long-Short-Term Memory, Rock. Mech. Rock. Eng., 57, 1353–1374, <a href="https://doi.org/10.1007/s00603-023-03600-z" target="_blank">https://doi.org/10.1007/s00603-023-03600-z</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>16</label><mixed-citation>
      
Fan, X., Cheng, J. Y., Wang, Y. H., Li, S., Duan, J. H., and Wang, P.: Intelligent recognition of coal mine microseismic signal based on wavelet scattering decomposition transform, J. China. Coal. Soc., 47, 2722–2731, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>17</label><mixed-citation>
      
Hochreiter, S. and Schmidhuber, J.: Long short-term memory, Neural. Comput., 9, 1735–1780, <a href="https://doi.org/10.1162/neco.1997.9.8.1735" target="_blank">https://doi.org/10.1162/neco.1997.9.8.1735</a>, 1997.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>18</label><mixed-citation>
      
Huang, W., Zou, X. Y., and Qiao, R.: Scattering wavelet based deep network for image classification, in: Communications in Computer and Information Science, edited by: Li, G., Filipe, J., and Xu, Z. W., Springer Singapore, 459–469, <a href="https://doi.org/10.1007/978-981-10-8530-7_45" target="_blank">https://doi.org/10.1007/978-981-10-8530-7_45</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>19</label><mixed-citation>
      
Khan, A. A., Dhawan, A., Akhlaghi, N., Majdi, J. A., and Sikdar, S.: Application of wavelet scattering networks in classification of ultrasound image sequences, in: Proceedings of the 2017 IEEE International Ultrasonics Symposium (IUS), Washington, DC, USA, 6–9 September 2017, 1–4, <a href="https://doi.org/10.1109/ULTSYM.2017.8091649" target="_blank">https://doi.org/10.1109/ULTSYM.2017.8091649</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>20</label><mixed-citation>
      
Koch, K. and Pilger, C.: Infrasound observations from the site of past underground nuclear explosions in North Korea, Geophys. J. Int., 216, 182–200, <a href="https://doi.org/10.1093/gji/ggy381" target="_blank">https://doi.org/10.1093/gji/ggy381</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>21</label><mixed-citation>
      
Kogelnig, A., Hübl, J., Suriñach, E., Vilajosana, I., and McArdell, B. W.: Infrasound produced by debris flow: propagation and frequency content evolution, Nat. Hazard., 70, 1713–1733, <a href="https://doi.org/10.1007/s11069-011-9741-8" target="_blank">https://doi.org/10.1007/s11069-011-9741-8</a>, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>22</label><mixed-citation>
      
Krizhevsky, A., Sutskever, I., and Hinton, G. E.: ImageNet classification with deep convolutional neural networks, J. Commun. ACM, 60, 84–90, <a href="https://doi.org/10.1145/3065386" target="_blank">https://doi.org/10.1145/3065386</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>23</label><mixed-citation>
      
Lee, S. and Hong, S. G.: Classification of nuclear activity types for neighboring countries of South Korea using machine learning techniques with xenon isotopic activity ratios, Nucl. Eng. Technol., 56, 1372–1384, <a href="https://doi.org/10.1016/j.net.2023.11.042" target="_blank">https://doi.org/10.1016/j.net.2023.11.042</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>24</label><mixed-citation>
      
Leng X. P., Feng, L. Y., Ou, O., Du, X. L., Liu, D. L., and Tang, X.: Debris flow infrasound recognition method based on improved LeNet-5 network, Sustainability-Basel, 14, 15925, <a href="https://doi.org/10.3390/su142315925" target="_blank">https://doi.org/10.3390/su142315925</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>25</label><mixed-citation>
      
Li, H. R., Li, X. H., Tan, X. F., Liu, T. Y., Zhang, Y., Niu, C., and Liu, J. H.: Infrasound event classification fusion model based on multiscale SE-CNN and BiLSTM, Appl. Geophys., 21, 579–592, <a href="https://doi.org/10.1007/s11770-024-1089-4" target="_blank">https://doi.org/10.1007/s11770-024-1089-4</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>26</label><mixed-citation>
      
Li, H. R., Li, X. H., Liu, J. H., Wang, Y. T., Liu, Z. G., and Zeng, X. N.: Machine learning-driven classification of natural disasters via parallel confidence fusion, Mach. Learn.: Sci. Technol., 7, 015028, <a href="https://doi.org/10.1088/2632-2153/ae3c58" target="_blank">https://doi.org/10.1088/2632-2153/ae3c58</a>, 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>27</label><mixed-citation>
      
Li, J. H., Ke, L., Du, Q., Ding, X. D., Chen, X. M., and Wang, D. N.: Heart sound signal classification algorithm: a combination of wavelet scattering transform and twin Support Vector Machine, IEEE Access, 7, 179339–179348, <a href="https://doi.org/10.1109/access.2019.2959081" target="_blank">https://doi.org/10.1109/access.2019.2959081</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>28</label><mixed-citation>
      
Li, M., Liu, X., and Liu, X.: Infrasound signal classification based on spectral entropy and support vector machine, Appl. Acoust., 113, 116–120, <a href="https://doi.org/10.1016/j.apacoust.2016.06.019" target="_blank">https://doi.org/10.1016/j.apacoust.2016.06.019</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>29</label><mixed-citation>
      
Lilly, J. M. and Olhede, S. C.: On the analytic wavelet transform, IEEE Trans. Inf. Theory., 56, 4135–4156, <a href="https://doi.org/10.1109/tit.2010.2050935" target="_blank">https://doi.org/10.1109/tit.2010.2050935</a>, 2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>30</label><mixed-citation>
      
Liu, L., Wu, J., Li, D., Senhadji, L., and Shu, H.: Fractional wavelet scattering network and applications, IEEE Trans. Biomed. Eng., 66, 553–563, <a href="https://doi.org/10.1109/tbme.2018.2850356" target="_blank">https://doi.org/10.1109/tbme.2018.2850356</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>31</label><mixed-citation>
      
Liu, X., Li, M., Tang, W., Wang, S., and Wu, X.: A new classification method of infrasound events using Hilbert-Huang Transform and Support Vector Machine, Math. Probl. Eng., 2014, 1–6, <a href="https://doi.org/10.1155/2014/456818" target="_blank">https://doi.org/10.1155/2014/456818</a>, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>32</label><mixed-citation>
      
Liu, Y. X., Zhang, B. Q., Kong, F. T., Wang, B., Luo, C. M., and Ma, L.: Underwater acoustic classification using wavelet scattering transform and convolutional neural network with limited dataset, Appl. Acoust., 232, 110564, <a href="https://doi.org/10.1016/j.apacoust.2025.110564" target="_blank">https://doi.org/10.1016/j.apacoust.2025.110564</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>33</label><mixed-citation>
      
Lone, A. W. and Aydin, N.: Wavelet scattering transform based doppler signal classification, Comput. Biol. Med., 167, 107611, <a href="https://doi.org/10.1016/j.compbiomed.2023.107611" target="_blank">https://doi.org/10.1016/j.compbiomed.2023.107611</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>34</label><mixed-citation>
      
Lu, X., Liu, S., Chen, X., and Zhang, K.: Identification of the lubrication state of journal bearings based on acoustic emission and WST-CNN collaboration, J. Vib. Shock., 42, 71–77, <a href="https://doi.org/10.13465/j.cnki.jvs.2023.22.008" target="_blank">https://doi.org/10.13465/j.cnki.jvs.2023.22.008</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>35</label><mixed-citation>
      
Luong, T., Pham, H., and Manning, C. D.: Effective approaches to attention-based neural machine translation, in: Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP), Lisbon, Portugal, 17 August 2015, 1412–1421, <a href="https://doi.org/10.18653/v1/D15-1166" target="_blank">https://doi.org/10.18653/v1/D15-1166</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>36</label><mixed-citation>
      
Marchetti, E., Walter, F., Barfucci, G., Genco, R., Wenner, M., Ripepe, M., McArdell, B., and Price, C.: Infrasound array analysis of debris flow activity and implication for early warning, J. Geophys. Res.-Earth, 124, 567–587, <a href="https://doi.org/10.1029/2018jf004785" target="_blank">https://doi.org/10.1029/2018jf004785</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>37</label><mixed-citation>
      
Park, J., Che, I. Y., Stump, B., Hayward, C., Dannemann, F., SeongJu, J., Kwong, K., McComas, S., Oldham, H. R., Scales, M. M., and Wright, V.: Characteristics of infrasound signals from North Korean underground nuclear explosions on 2016 January 6 and September 9, Geophys. J. Int., 214, 1865–1885, <a href="https://doi.org/10.1093/gji/ggy252" target="_blank">https://doi.org/10.1093/gji/ggy252</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>38</label><mixed-citation>
      
Pásztor, M., Czanik, C., and Bondár, I.: A single array approach for infrasound signal discrimination from quarry blasts via machine learning, Remote. Sens., 15, 1657, <a href="https://doi.org/10.3390/rs15061657" target="_blank">https://doi.org/10.3390/rs15061657</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>39</label><mixed-citation>
      
Priya, B. L., Jayalakshmy, S., Pragatheeswaran, J. K., Saraswathi, D., and Poonguzhali, N.: Scattering convolutional network based predictive model for cognitive activity of brain using empirical wavelet decomposition, Biomed. Signal. Process. Control., 66, 102501, <a href="https://doi.org/10.1016/j.bspc.2021.102501" target="_blank">https://doi.org/10.1016/j.bspc.2021.102501</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>40</label><mixed-citation>
      
Shirodkar, V. R., Edla, D. R., and Kumari, A.: Advancing motor imagery EEG classification through wavelet scattering transforms and 1D transformers, Procedia Comput. Sci., 258, 2860–2869, <a href="https://doi.org/10.1016/j.procs.2025.04.546" target="_blank">https://doi.org/10.1016/j.procs.2025.04.546</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>41</label><mixed-citation>
      
Singhal, S. and Kumar, M.: SPTDMD-WST: Arrhythmia classification from spatiotemporal modes of dynamic mode decomposition using wavelet scattering transform, Biomed. Signal. Process. Control., 92, 105983, <a href="https://doi.org/10.1016/j.bspc.2024.105983" target="_blank">https://doi.org/10.1016/j.bspc.2024.105983</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>42</label><mixed-citation>
      
Stéphane, M.: CHAPTER 7-Wavelet Bases, in: A Wavelet Tour of Signal Processing, Academic Press, 263–376, <a href="https://doi.org/10.1016/B978-0-12-374370-1.00011-2" target="_blank">https://doi.org/10.1016/B978-0-12-374370-1.00011-2</a>, 2009.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>43</label><mixed-citation>
      
Souli, S., Amami, R., and Yahia, S. B.: A robust pathological voices recognition system based on DCNN and scattering transform, Appl. Acoust., 177, 107854, <a href="https://doi.org/10.1016/j.apacoust.2020.107854" target="_blank">https://doi.org/10.1016/j.apacoust.2020.107854</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>44</label><mixed-citation>
      
Tan, X. F., Li, X. H., Liu, J. H., Li, G., and Yu, X. T.: Classification of chemical explosion and earthquake infrasound based on 1-D convolutional neural network, J. Appl. Acoust., 40, 457–467, <a href="https://doi.org/10.11684/j.issn.1000-310X.2021.03.018" target="_blank">https://doi.org/10.11684/j.issn.1000-310X.2021.03.018</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>45</label><mixed-citation>
      
Tan, X. F., Li, X. H., Niu, C., Zeng, X. N., Li, H, R., and Liu, T. Y.: Classification method of infrasound events based on the MVIDA algorithm and MS-SE-ResNet, Appl. Geophys., 21, 667–679, <a href="https://doi.org/10.1007/s11770-024-1112-9" target="_blank">https://doi.org/10.1007/s11770-024-1112-9</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>46</label><mixed-citation>
      
Thüring, T., Schoch, M., Herwijnen, A., and Schweizer, J.: Robust snow avalanche detection using supervised machine learning with infrasonic sensor arrays, Cold. Reg. Sci. Technol., 111, 60–66, <a href="https://doi.org/10.1016/j.coldregions.2014.12.014" target="_blank">https://doi.org/10.1016/j.coldregions.2014.12.014</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>47</label><mixed-citation>
      
Toney, L., Fee, D., Witsil, A., and Matoza, R. S.: Waveform features strongly control subcrater classification performance for a large labeled volcano infrasound dataset, Seism. Rec., 2, 167–175, <a href="https://doi.org/10.1785/0320220019" target="_blank">https://doi.org/10.1785/0320220019</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>48</label><mixed-citation>
      
Tuan, D. P.: Classification of motor-imagery tasks using a large eeg dataset by fusing classifiers learning on wavelet-scattering features, IEEE Trans. Neural. Syst. Rehabil. Eng., 31, 1097–1107, <a href="https://doi.org/10.1109/TNSRE.2023.3241241" target="_blank">https://doi.org/10.1109/TNSRE.2023.3241241</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>49</label><mixed-citation>
      
Wang, H., Li, S., Zhou, Y., and Chen, S.: SAR automatic target recognition using a roto-translational invariant wavelet-scattering convolution network, Remote Sens., 10, 501, <a href="https://doi.org/10.3390/rs10040501" target="_blank">https://doi.org/10.3390/rs10040501</a>, 2018a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib50"><label>50</label><mixed-citation>
      
Wang, L., Liu, B., Wang, H., Zhong, G., and Dong, J.: Deep gabor scattering network for image classification, in: Pattern Recognition and Computer Vision, 2 November 2018, Springer International Publishing, 332–343, <a href="https://doi.org/10.1007/978-3-030-03335-4_29" target="_blank">https://doi.org/10.1007/978-3-030-03335-4_29</a>, 2018b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib51"><label>51</label><mixed-citation>
      
Wang, W., Gao, X., and Le, L.: Study of the similarities in scale models of a single-layer spherical lattice shell structure under the effect of internal explosion, Shock. Vib., 4, 1–13, <a href="https://doi.org/10.1155/2017/9181729" target="_blank">https://doi.org/10.1155/2017/9181729</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib52"><label>52</label><mixed-citation>
      
Wang, Z. X., Gu, W. B., Liang, T., Zhao, S. T., Chen, P., and Yu, L. F.: Monitoring and prediction of the vibration intensity of seismic waves induced in underwater rock by underwater drilling and blasting, Def. Technol., 18, 109–118, <a href="https://doi.org/10.1016/j.dt.2020.10.007" target="_blank">https://doi.org/10.1016/j.dt.2020.10.007</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib53"><label>53</label><mixed-citation>
      
Wen, Y. D., Ren, W. T., Yang, Y., Chen, J., Han, Y., and Ye, Y. Z.: Analysis of infrasound signals generated by CZ-7 teleport 2 launch vehicle, J. Ordnance Equip. Eng., 40, 70–73, <a href="https://doi.org/10.11809/bqzbgcxb2019.09.015" target="_blank">https://doi.org/10.11809/bqzbgcxb2019.09.015</a>, 2019.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib54"><label>54</label><mixed-citation>
      
Wiatowski, T. and Bolcskei, H.: A mathematical theory of deep convolutional neural networks for feature extraction, IEEE Trans. Inf. Theory., 64, 1845–1866, <a href="https://doi.org/10.1109/TIT.2017.2776228" target="_blank">https://doi.org/10.1109/TIT.2017.2776228</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib55"><label>55</label><mixed-citation>
      
Yin, H., Lu, Y., Wu, Y. H., Cheng, W., Pang, X. L., and Li, P.: Infrasound signal classification fusion model based on double-branch and multi-scale CNN and LSTM, Acoustics, 8, 21, <a href="https://doi.org/10.3390/acoustics8020021" target="_blank">https://doi.org/10.3390/acoustics8020021</a>, 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib56"><label>56</label><mixed-citation>
      
Yuan, L., Liu, D. L., Sang, X. J., Zhang, S. J., and Chen, Q.: Debris flow infrasound signal recognition approach based on improved AlexNet, Computer and Modernization, 3, 1–6, <a href="https://doi.org/10.3969/j.issn.1006-2475.2024.03.001" target="_blank">https://doi.org/10.3969/j.issn.1006-2475.2024.03.001</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib57"><label>57</label><mixed-citation>
      
Zhang, A., Chun, S., Cheng, Z., and Zhao, P.: Predicting the core thermal hydraulic parameters with a gated recurrent unit model based on the soft attention mechanism, Nucl. Eng. Technol., 56, 2343–2351, <a href="https://doi.org/10.1016/j.net.2024.01.045" target="_blank">https://doi.org/10.1016/j.net.2024.01.045</a>, 2024a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib58"><label>58</label><mixed-citation>
      
Zhang, H. Y., Zhao, Z. J., Liu, C., Duan, M., Lu, Z. G., and Wang, H.: Classification of motor imagery EEG signals using wavelet scattering transform and Bi-directional long short-term memory networks, Biocybern. Biomed. Eng., 44, 874–884, <a href="https://doi.org/10.1016/j.bbe.2024.11.003" target="_blank">https://doi.org/10.1016/j.bbe.2024.11.003</a>, 2024b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib59"><label>59</label><mixed-citation>
      
Zhang, Q., Wang, P., Pedrycz, W., and Li, Z.: Neighborhood entropy guided by a decision attribute and its applications in multi-source information fusion and attribute selection, Appl. Soft. Comput., 167, 112380, <a href="https://doi.org/10.1016/j.asoc.2024.112380" target="_blank">https://doi.org/10.1016/j.asoc.2024.112380</a>, 2024c.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib60"><label>60</label><mixed-citation>
      
Zhang, X., Shen, J., Li, J., and Liu, X.: An instance-oriented multi-source information fusion technique based on neighborhood granules, Appl. Soft. Comput., 181, 113483, <a href="https://doi.org/10.1016/j.asoc.2025.113483" target="_blank">https://doi.org/10.1016/j.asoc.2025.113483</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib61"><label>61</label><mixed-citation>
      
Zhao, Z. J., Chen, W., Ji, P. F., Teng, P. X., Lyu, J., and Yang, J.: A method for classification of few-shot infrasound signals applying prototype network, J. Appl. Acoust., 43, 1193–1202, 2024.

    </mixed-citation></ref-html>--></article>
