<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">NHESS</journal-id><journal-title-group>
    <journal-title>Natural Hazards and Earth System Sciences</journal-title>
    <abbrev-journal-title abbrev-type="publisher">NHESS</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Nat. Hazards Earth Syst. Sci.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1684-9981</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/nhess-23-3199-2023</article-id><title-group><article-title>Testing machine learning models for heuristic building damage assessment
applied to the Italian Database<?xmltex \hack{\break}?> of Observed Damage (DaDO)</article-title><alt-title>Testing machine learning models</alt-title>
      </title-group><?xmltex \runningtitle{Testing machine learning models}?><?xmltex \runningauthor{S. Ghimire et al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Ghimire</surname><given-names>Subash</given-names></name>
          <email>subash.ghimire@univ-grenoble-alpes.fr</email>
        <ext-link>https://orcid.org/0000-0001-8119-7316</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Guéguen</surname><given-names>Philippe</given-names></name>
          
        <ext-link>https://orcid.org/0000-0001-6362-0694</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Pothon</surname><given-names>Adrien</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff3">
          <name><surname>Schorlemmer</surname><given-names>Danijel</given-names></name>
          
        </contrib>
        <aff id="aff1"><label>1</label><institution>ISTerre, Université Grenoble Alpes/CNRS/IRD/Université Gustave
Eiffel, Grenoble, <?xmltex \hack{\break}?>CS40700 38058 Grenoble CEDEX 9, France</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>AXA Group Risk Management, GIE AXA, 21 Avenue Matignon, 75008 Paris,
France</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>German Research Centre for Geosciences, Telegrafenberg, 14473 Potsdam,
Germany</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Subash Ghimire (subash.ghimire@univ-grenoble-alpes.fr)</corresp></author-notes><pub-date><day>5</day><month>October</month><year>2023</year></pub-date>
      
      <volume>23</volume>
      <issue>10</issue>
      <fpage>3199</fpage><lpage>3218</lpage>
      <history>
        <date date-type="received"><day>20</day><month>January</month><year>2023</year></date>
           <date date-type="rev-request"><day>7</day><month>February</month><year>2023</year></date>
           <date date-type="rev-recd"><day>27</day><month>June</month><year>2023</year></date>
           <date date-type="accepted"><day>23</day><month>August</month><year>2023</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2023 Subash Ghimire et al.</copyright-statement>
        <copyright-year>2023</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://nhess.copernicus.org/articles/23/3199/2023/nhess-23-3199-2023.html">This article is available from https://nhess.copernicus.org/articles/23/3199/2023/nhess-23-3199-2023.html</self-uri><self-uri xlink:href="https://nhess.copernicus.org/articles/23/3199/2023/nhess-23-3199-2023.pdf">The full text article is available as a PDF file from https://nhess.copernicus.org/articles/23/3199/2023/nhess-23-3199-2023.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d1e125">Assessing or forecasting seismic damage to buildings is an essential issue
for earthquake disaster management. In this study, we explore the efficacy
of several machine learning models for damage characterization, trained and
tested on the database of damage observed after Italian earthquakes (the Database of Observed Damage – DaDO).
Six models were considered: regression- and classification-based machine
learning models, each using random forest, gradient boosting, and extreme
gradient boosting. The structural features considered were divided into two
groups: all structural features provided by DaDO or only those considered to
be the most reliable and easiest to collect (age, number of storeys, floor
area, building height). Macroseismic intensity was also included as an input
feature. The seismic damage per building was determined according to the
EMS-98 scale observed after seven significant earthquakes occurring in
several Italian regions. The results showed that extreme gradient boosting
classification is statistically the most efficient method, particularly when
considering the basic structural features and grouping the damage according
to the traffic-light-based system used; for example, during the
post-disaster period (green, yellow, and red), 68 % of buildings were
correctly classified. The results obtained by the machine-learning-based
heuristic model for damage assessment are of the same order of accuracy
(error values were less than 17 %) as those obtained by the traditional
RISK-UE method. Finally, the machine learning analysis found that the
importance of structural features with respect to damage was conditioned by
the level of damage considered.</p>
  </abstract>
    
<funding-group>
<award-group id="gs1">
<funding-source>H2020 Marie Skłodowska-Curie Actions</funding-source>
<award-id>813137</award-id>
</award-group>
<award-group id="gs2">
<funding-source>Agence Nationale de la Recherche</funding-source>
<award-id>ANR10LABX56</award-id>
</award-group>
<award-group id="gs3">
<funding-source>AXA Research Fund</funding-source>
<award-id>New Probabilistic Assessment of Seismic Hazard, Losses and Risks in Strong Seismic Prone Regions</award-id>
</award-group>
</funding-group>
</article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d1e137">Population growth worldwide increases exposure to natural hazards,
increasing consequences in terms of global economic and human losses. For
example, between 1985 and 2014, the world's population increased by 50 %
and average annual losses due to natural disasters increased from USD 14 billion to over USD 140 billion (Silva et al., 2019). Among other natural
hazards, earthquakes represent one-fifth of total annual economic losses and
cause more than 20 000 deaths per year
(Daniell et al., 2017; Silva et
al., 2019). To develop effective seismic risk reduction policies,
decision-makers and stakeholders rely on a representation of consequences
when earthquakes affect the built environment. Two main risk metrics
generally considered at the global scale are associated with building
damage: direct economic losses due to costs of repair/replacement and loss
of life of inhabitants due to building damage. The damage is estimated by
combining the seismic hazard, exposure models, and vulnerability/fragility
functions (Silva et al., 2019).</p>
      <p id="d1e140">For scenario-based risk assessment, damage and related consequences are
computed for a single earthquake defined in terms of magnitude, location,
and other seismological features. Many methods have been developed to
characterize the<?pagebreak page3200?> urban environment for exposure models. In particular,
damage assessment requires vulnerability/fragility functions for all types
of existing buildings, defined according to their design characteristics
(shape, position, materials, height, etc.) and grouped in a building
taxonomy (e.g. among other conventional methods, FEMA,
2003; Grünthal, 1998; Guéguen et al., 2007; Lagomarsino and
Giovinazzi, 2006; Mouroux and Le Brun, 2006;
Silva et al., 2022). At the regional/country
scale, damage assessment is therefore confronted with the difficulty of
accurately characterizing exposure according to the required criteria and
assigning appropriate vulnerability/fragility functions to building
features. Unfortunately, the necessary information is often sparse and
incomplete, and the exposure model development suffers from economic and
time constraints.</p>
      <p id="d1e143">Over the past decade, there has been growing interest in artificial
intelligence methods for seismic risk assessment due to their superior
computational efficiency, their easy handling of complex problems, and the
incorporation of uncertainties (e.g.
Riedel et al., 2014, 2015; Azimi et al., 2020; Ghimire et al., 2022; Hegde
and Rokseth, 2020; Kim et al., 2020; Mangalathu and Jeon, 2020; Morfidis
and Kostinakis, 2018; Salehi and Burgueño, 2018; Seo et al., 2012; Sun
et al., 2021; Wang et al., 2021; Xie et al., 2020; Y. Xu et al., 2020; Z. Xu
et al., 2020). In particular, several studies have tested the effectiveness
of machine learning methods in associating damage degrees with basic
building features and spatially distributed seismic demand with acceptable
accuracy compared with conventional methods or with post-earthquake
observations (e.g.
Riedel et al., 2014, 2015; Guettiche et al., 2017; Harirchian et al., 2021;
Mangalathu et al., 2020; Roeslin et al., 2020; Stojadinović et al.,
2021; Ghimire et al., 2022). In parallel, significant efforts have been made
to collect post-earthquake building damage observations after damaging
earthquakes (Dolce
et al., 2019; MINVU, 2010; MTPTC, 2010; NPC, 2015). With more than 10 000
samples compiled, the Database of Observed Damage (DaDO) in Italy, a
platform of the Civil Protection Department, developed by the Eucentre
Foundation (Dolce et al., 2019), allows exploration of the value of
heuristic vulnerability functions calibrated on observations
(Lagomarsino et al.,
2021), as well as the training of heuristic functions using machine learning
models (Ghimire et al., 2022) and
considering sparse and incomplete building features.</p>
      <p id="d1e146">The main objective of this study is to investigate the effectiveness of
several machine learning models trained and tested on information from
DaDO to develop a heuristic model for damage assessment. The model may be
classified as heuristic because it applies a problem-solving approach in
which a calculated guess based on previous experience is considered for
damage assessment (as opposed to applying algorithms that effectively
eliminate the approximation). The damage is thus estimated in a non-rigorous
way defined during the training phase, and the results must be validated and
then tested against observed damage. By analogy with psychology, this
procedure can reduce the cognitive load associated with uncertainties when
making decisions based on damage assessment by explicitly considering the
uncertainties in the assessment, being aware of the incompleteness of the
information and the accuracy level to make a decision. The dataset and
methods are described in the Data and Method sections, respectively. The
fourth section presents the results of damage prediction produced by machine
learning models compared with conventional methods, followed by the Discussion
and Conclusions sections.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Data</title>
      <p id="d1e157">The Database of Observed Damage (DaDO; Dolce et al., 2019) is accessible
through a web-based geographic information system (GIS) platform and is designed to collect and share information
about building features, seismic ground motions, and observed damage
following major earthquakes in Italy from 1976 to 2019 (with the exclusion
of the 2016–2017 central Italy earthquake for which data processing is
ongoing). A framework was adopted to homogenize the different forms of
information collected and to translate the damage information into the
EMS-98 scale (Grünthal, 1998) using the method proposed by Dolce et
al. (2019). For this study, we selected building damage data from seven
earthquakes summarized in Table 1 and presented in Fig. 1.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1" specific-use="star"><?xmltex \currentcnt{1}?><label>Table 1</label><caption><p id="d1e163">Building damage data from DaDO for the seven earthquakes
considered in this study. “Ref” is the reference to the earthquake used in
the paper. “DL” is the number of the damage grade available in DaDO.
“NB” is the number of buildings considered in this study. AeDES is the
post-earthquake damage survey form, first introduced in 1997 and which became the official operational tool recognized by the Italian Civil
Protection Department in 2002.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="9">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="left"/>
     <oasis:colspec colnum="8" colname="col8" align="right"/>
     <oasis:colspec colnum="9" colname="col9" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Ref</oasis:entry>
         <oasis:entry colname="col2">Earthquake</oasis:entry>
         <oasis:entry colname="col3">Event date</oasis:entry>
         <oasis:entry colname="col4">Mag. (<inline-formula><mml:math id="M1" display="inline"><mml:mrow><mml:msub><mml:mi>M</mml:mi><mml:mi mathvariant="normal">w</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry rowsep="1" namest="col5" nameend="col6" align="center">Epicentre </oasis:entry>
         <oasis:entry colname="col7">Damage survey form</oasis:entry>
         <oasis:entry colname="col8">DL</oasis:entry>
         <oasis:entry colname="col9">NB</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5">Lat</oasis:entry>
         <oasis:entry colname="col6">Long</oasis:entry>
         <oasis:entry colname="col7"/>
         <oasis:entry colname="col8"/>
         <oasis:entry colname="col9"/>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">E1</oasis:entry>
         <oasis:entry colname="col2">Irpinia 1980</oasis:entry>
         <oasis:entry colname="col3">23 Nov 1980</oasis:entry>
         <oasis:entry colname="col4">6.9</oasis:entry>
         <oasis:entry colname="col5">40.91</oasis:entry>
         <oasis:entry colname="col6">15.37</oasis:entry>
         <oasis:entry colname="col7">Irpinia 1980</oasis:entry>
         <oasis:entry colname="col8">8</oasis:entry>
         <oasis:entry colname="col9">37 828</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">E2</oasis:entry>
         <oasis:entry colname="col2">Pollino 1998</oasis:entry>
         <oasis:entry colname="col3">9 Sep 1998</oasis:entry>
         <oasis:entry colname="col4">5.6</oasis:entry>
         <oasis:entry colname="col5">40.04</oasis:entry>
         <oasis:entry colname="col6">15.98</oasis:entry>
         <oasis:entry colname="col7">AeDES-1998</oasis:entry>
         <oasis:entry colname="col8">4</oasis:entry>
         <oasis:entry colname="col9">9485</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">E3</oasis:entry>
         <oasis:entry colname="col2">Molise–Puglia 2002</oasis:entry>
         <oasis:entry colname="col3">31 Oct 2002</oasis:entry>
         <oasis:entry colname="col4">5.9</oasis:entry>
         <oasis:entry colname="col5">41.79</oasis:entry>
         <oasis:entry colname="col6">14.87</oasis:entry>
         <oasis:entry colname="col7">AeDES-2000</oasis:entry>
         <oasis:entry colname="col8">4</oasis:entry>
         <oasis:entry colname="col9">6396</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">E4</oasis:entry>
         <oasis:entry colname="col2">Emilia-Romagna 2003</oasis:entry>
         <oasis:entry colname="col3">14 Sep 2003</oasis:entry>
         <oasis:entry colname="col4">5.3</oasis:entry>
         <oasis:entry colname="col5">44.33</oasis:entry>
         <oasis:entry colname="col6">11.45</oasis:entry>
         <oasis:entry colname="col7">AeDES-2000</oasis:entry>
         <oasis:entry colname="col8">4</oasis:entry>
         <oasis:entry colname="col9">239</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">E5</oasis:entry>
         <oasis:entry colname="col2">L'Aquila 2009</oasis:entry>
         <oasis:entry colname="col3">6 Apr 2009</oasis:entry>
         <oasis:entry colname="col4">6.3</oasis:entry>
         <oasis:entry colname="col5">42.34</oasis:entry>
         <oasis:entry colname="col6">13.34</oasis:entry>
         <oasis:entry colname="col7">AeDES-2008</oasis:entry>
         <oasis:entry colname="col8">4</oasis:entry>
         <oasis:entry colname="col9">37 999</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">E6</oasis:entry>
         <oasis:entry colname="col2">Emilia-Romagna 2012</oasis:entry>
         <oasis:entry colname="col3">20 May 2012</oasis:entry>
         <oasis:entry colname="col4">6.1</oasis:entry>
         <oasis:entry colname="col5">44.89</oasis:entry>
         <oasis:entry colname="col6">11.23</oasis:entry>
         <oasis:entry colname="col7">AeDES-2008</oasis:entry>
         <oasis:entry colname="col8">4</oasis:entry>
         <oasis:entry colname="col9">10 581</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">E7</oasis:entry>
         <oasis:entry colname="col2">Garfagnana–Lunigiana 2013</oasis:entry>
         <oasis:entry colname="col3">21 Jun 2013</oasis:entry>
         <oasis:entry colname="col4">5.3</oasis:entry>
         <oasis:entry colname="col5">44.15</oasis:entry>
         <oasis:entry colname="col6">10.14</oasis:entry>
         <oasis:entry colname="col7">AeDES-2008</oasis:entry>
         <oasis:entry colname="col8">4</oasis:entry>
         <oasis:entry colname="col9">1474</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table><?xmltex \gdef\@currentlabel{1}?></table-wrap>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1"><?xmltex \currentcnt{1}?><?xmltex \def\figurename{Figure}?><label>Figure 1</label><caption><p id="d1e480">Geographic location of the buildings considered in this study.</p></caption>
        <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://nhess.copernicus.org/articles/23/3199/2023/nhess-23-3199-2023-f01.png"/>

      </fig>

      <p id="d1e490">The converted EMS-98 damage grade (DG) ranges from damage grade DG0 (no
damage) to DG5 (total collapse). The building features are available for
each individual building and relate to the shape and design of the building
and the built-up environment (Table 2, Fig. 2) as follows:
<list list-type="bullet"><list-item>
      <?pagebreak page3201?><p id="d1e495">building location – defined by its latitude
and longitude, assigned using either the exact address of the building if
available or the address of the local administrative centre
(Dolce et al., 2019);</p></list-item><list-item>
      <p id="d1e499">number of storeys – total number of floors above the surface of the ground;</p></list-item><list-item>
      <p id="d1e503">age of building – time difference between the date of the earthquake and the
date of building construction/renovation;</p></list-item><list-item>
      <p id="d1e507">height of building – total height of the building above the surface of the
ground, in metres;</p></list-item><list-item>
      <p id="d1e511">floor area – average of the storey surface area, in square metres;</p></list-item><list-item>
      <p id="d1e515">ground slope condition – four types of ground slope conditions (flat, mild slope, steep slope, and ridge);</p></list-item><list-item>
      <p id="d1e519">roof type – four types of roofs (thrusting heavy roof,
non-thrusting heavy roof, thrusting light roof, and non-thrusting light
roof);</p></list-item><list-item>
      <p id="d1e523">position of building – indication of the building's position in the block (isolated, extreme, corner, and intermediate);</p></list-item><list-item>
      <p id="d1e527">regularity – building regularity in terms of plan and elevation, classified
as either irregular or regular;</p></list-item><list-item>
      <p id="d1e531">construction material – vertical elements of good- and poor-quality masonry,
good- and poor-quality mixed frame masonry, reinforced concrete frame and
wall, steel frame, and other.</p></list-item></list>
For features defined as value ranges (e.g. date of construction/renovation,
floor area, and building height), the average value was used. Furthermore,
the Irpinia 1980 building damage portfolio (E1) was constructed using the
specific Irpinia 1980 damage survey form, while the AeDES damage survey form
was used for the others. The Irpinia 1980 dataset will therefore be analysed
separately.</p>

<?xmltex \floatpos{p}?><table-wrap id="Ch1.T2" specific-use="star"><?xmltex \currentcnt{2}?><label>Table 2</label><caption><p id="d1e538">Distribution of the different features used in this study.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="7">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:colspec colnum="5" colname="col5" align="left"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="left"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">No.</oasis:entry>
         <oasis:entry namest="col2" nameend="col4" align="center">Parameters </oasis:entry>
         <oasis:entry colname="col5">Data type</oasis:entry>
         <oasis:entry colname="col6">Distribution</oasis:entry>
         <oasis:entry colname="col7">Remarks</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">(%)</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">1</oasis:entry>
         <oasis:entry colname="col2">Damage</oasis:entry>
         <oasis:entry colname="col3">No damage</oasis:entry>
         <oasis:entry colname="col4">DG0</oasis:entry>
         <oasis:entry colname="col5">Categorical</oasis:entry>
         <oasis:entry colname="col6">43.63</oasis:entry>
         <oasis:entry colname="col7">Fig. 2a</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">grades</oasis:entry>
         <oasis:entry colname="col3">Slight damage</oasis:entry>
         <oasis:entry colname="col4">DG1</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">28.90</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">(DGs)</oasis:entry>
         <oasis:entry colname="col3">Moderate damage</oasis:entry>
         <oasis:entry colname="col4">DG2</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">7.41</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Substantial damage</oasis:entry>
         <oasis:entry colname="col4">DG3</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">12.48</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Very heavy damage</oasis:entry>
         <oasis:entry colname="col4">DG4</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">3.94</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Total collapse</oasis:entry>
         <oasis:entry colname="col4">DG5</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">3.65</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">2</oasis:entry>
         <oasis:entry colname="col2">Number of</oasis:entry>
         <oasis:entry colname="col3">0–3</oasis:entry>
         <oasis:entry colname="col4">NF1</oasis:entry>
         <oasis:entry colname="col5">Numerical</oasis:entry>
         <oasis:entry colname="col6">85.81</oasis:entry>
         <oasis:entry colname="col7">Fig. 2b</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">storeys</oasis:entry>
         <oasis:entry colname="col3">3–5</oasis:entry>
         <oasis:entry colname="col4">NF2</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">13.01</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M2" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> 5</oasis:entry>
         <oasis:entry colname="col4">NF3</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">1.19</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">3</oasis:entry>
         <oasis:entry colname="col2">Age</oasis:entry>
         <oasis:entry colname="col3">0–20</oasis:entry>
         <oasis:entry colname="col4">AG1</oasis:entry>
         <oasis:entry colname="col5">Numerical</oasis:entry>
         <oasis:entry colname="col6">15.22</oasis:entry>
         <oasis:entry colname="col7">Fig. 2c</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">(years)</oasis:entry>
         <oasis:entry colname="col3">21–40</oasis:entry>
         <oasis:entry colname="col4">AG2</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">18.81</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">41–60</oasis:entry>
         <oasis:entry colname="col4">AG3</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">34.15</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">61–80</oasis:entry>
         <oasis:entry colname="col4">AG4</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">21.34</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M3" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> 80</oasis:entry>
         <oasis:entry colname="col4">AG5</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">10.49</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">4</oasis:entry>
         <oasis:entry colname="col2">Floor area</oasis:entry>
         <oasis:entry colname="col3">0–50</oasis:entry>
         <oasis:entry colname="col4">A1</oasis:entry>
         <oasis:entry colname="col5">Numerical</oasis:entry>
         <oasis:entry colname="col6">22.16</oasis:entry>
         <oasis:entry colname="col7">Fig. 2d</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">(square metres)</oasis:entry>
         <oasis:entry colname="col3">50–100</oasis:entry>
         <oasis:entry colname="col4">A2</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">34.73</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">100–150</oasis:entry>
         <oasis:entry colname="col4">A3</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">22.53</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">150–200</oasis:entry>
         <oasis:entry colname="col4">A4</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">8.32</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M4" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> 200</oasis:entry>
         <oasis:entry colname="col4">A5</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">12.26</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">5</oasis:entry>
         <oasis:entry colname="col2">Height</oasis:entry>
         <oasis:entry colname="col3">0–10</oasis:entry>
         <oasis:entry colname="col4">H1</oasis:entry>
         <oasis:entry colname="col5">Numerical</oasis:entry>
         <oasis:entry colname="col6">87.78</oasis:entry>
         <oasis:entry colname="col7">Fig. 2e</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">(metres)</oasis:entry>
         <oasis:entry colname="col3">10–15</oasis:entry>
         <oasis:entry colname="col4">H2</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">10.69</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M5" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> 15</oasis:entry>
         <oasis:entry colname="col4">H3</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">1.50</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">6</oasis:entry>
         <oasis:entry colname="col2">Position</oasis:entry>
         <oasis:entry colname="col3">Corner</oasis:entry>
         <oasis:entry colname="col4">P1</oasis:entry>
         <oasis:entry colname="col5">Categorical</oasis:entry>
         <oasis:entry colname="col6">9.71</oasis:entry>
         <oasis:entry colname="col7">Fig. 2f</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Extreme</oasis:entry>
         <oasis:entry colname="col4">P2</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">24.47</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Internal</oasis:entry>
         <oasis:entry colname="col4">P3</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">22.80</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Isolated</oasis:entry>
         <oasis:entry colname="col4">P4</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">43.02</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">7</oasis:entry>
         <oasis:entry colname="col2">Ground</oasis:entry>
         <oasis:entry colname="col3">Ridge</oasis:entry>
         <oasis:entry colname="col4">GS1</oasis:entry>
         <oasis:entry colname="col5">Categorical</oasis:entry>
         <oasis:entry colname="col6">2.62</oasis:entry>
         <oasis:entry colname="col7">Fig. 2g</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">slope</oasis:entry>
         <oasis:entry colname="col3">Plain</oasis:entry>
         <oasis:entry colname="col4">GS2</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">34.25</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Moderate slope</oasis:entry>
         <oasis:entry colname="col4">GS3</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">43.74</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Steep slope</oasis:entry>
         <oasis:entry colname="col4">GS4</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">20.39</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">8</oasis:entry>
         <oasis:entry colname="col2">Regularity</oasis:entry>
         <oasis:entry colname="col3">Irregular in plan and elevation</oasis:entry>
         <oasis:entry colname="col4">IRe</oasis:entry>
         <oasis:entry colname="col5">Categorical</oasis:entry>
         <oasis:entry colname="col6">22.28</oasis:entry>
         <oasis:entry colname="col7">Fig. 2h</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Regular in plan and elevation</oasis:entry>
         <oasis:entry colname="col4">Re</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">77.72</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">9</oasis:entry>
         <oasis:entry colname="col2">Roof</oasis:entry>
         <oasis:entry colname="col3">Heavy, no thrust</oasis:entry>
         <oasis:entry colname="col4">R1</oasis:entry>
         <oasis:entry colname="col5">Categorical</oasis:entry>
         <oasis:entry colname="col6">36.43</oasis:entry>
         <oasis:entry colname="col7">Fig. 2i</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">type</oasis:entry>
         <oasis:entry colname="col3">Heavy thrust</oasis:entry>
         <oasis:entry colname="col4">R2</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">11.25</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Light thrust</oasis:entry>
         <oasis:entry colname="col4">R3</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">26.48</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Light, no thrust</oasis:entry>
         <oasis:entry colname="col4">R4</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">25.83</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">10</oasis:entry>
         <oasis:entry colname="col2">Material</oasis:entry>
         <oasis:entry colname="col3">Masonry, poor quality</oasis:entry>
         <oasis:entry colname="col4">CM1</oasis:entry>
         <oasis:entry colname="col5">Categorical</oasis:entry>
         <oasis:entry colname="col6">36.51</oasis:entry>
         <oasis:entry colname="col7">Fig. 2j</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Masonry, good quality</oasis:entry>
         <oasis:entry colname="col4">CM2</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">28.96</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Mixed frame masonry, poor quality</oasis:entry>
         <oasis:entry colname="col4">CM3</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">2.64</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Mixed frame masonry, good quality</oasis:entry>
         <oasis:entry colname="col4">CM4</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">5.21</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Reinforced concrete frame</oasis:entry>
         <oasis:entry colname="col4">CM5</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">21.31</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Reinforced concrete wall</oasis:entry>
         <oasis:entry colname="col4">CM6</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">0.42</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Steel frame</oasis:entry>
         <oasis:entry colname="col4">CM7</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">0.09</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">Other</oasis:entry>
         <oasis:entry colname="col4">CM8</oasis:entry>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">4.10</oasis:entry>
         <oasis:entry colname="col7"/>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table><?xmltex \gdef\@currentlabel{2}?></table-wrap>

      <p id="d1e1607">Building damage data from earthquake surveys other than the Irpinia 1980
earthquake damage survey primarily include damaged buildings. This is
because the data were collected based on requests for damage assessments
after the earthquake event (Dolce et al., 2019). The damage information in
the DaDO database is still relevant for testing the machine learning models
for heuristic damage assessment. Mixing these datasets to train machine
learning models can lead to biased outcomes. Therefore, the machine learning
models were developed on the other earthquake dataset excluding the Irpinia
dataset, and the Irpinia earthquake dataset was used only in the testing
phase.</p>
      <p id="d1e1610">The distribution of the samples is very imbalanced (Fig. 2): for example,
there is a small proportion of buildings in the DG4 <inline-formula><mml:math id="M6" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> DG5 categories (7.59 %) and a
large majority of masonry (65.47 %) compared to reinforced concrete frame
(21.31 %) buildings. This imbalance should be taken into account when
defining the machine learning models.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2" specific-use="star"><?xmltex \currentcnt{2}?><?xmltex \def\figurename{Figure}?><label>Figure 2</label><caption><p id="d1e1622">Distribution of the different features in the database. E1, E2,
E3, E4, E5, E6, and E7 represent Irpinia 1980, Pollino 1998,
Molise–Puglia 2002, Emilia-Romagna 2003, L'Aquila 2009, Emilia-Romagna 2012,
and Garfagnana–Lunigiana 2013 building damage portfolios, respectively. The
<inline-formula><mml:math id="M7" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> axis is the percentage distribution. and the <inline-formula><mml:math id="M8" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> axis is <bold>(a)</bold> damage grade,
<bold>(b)</bold> number of storeys (NF1: 0–3, NF2: 3–5, NF3: <inline-formula><mml:math id="M9" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> 5 storeys), <bold>(c)</bold>
building age (AG1: 0–20, AG2: 21–40, AG3: 41–60, AG4: 61–80, AG5:
<inline-formula><mml:math id="M10" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> 80 years), <bold>(d)</bold> floor area (A1: 0–50, A2: 51–100, A3: 101–150, A4:
151–200, A5: <inline-formula><mml:math id="M11" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> 200 m<inline-formula><mml:math id="M12" display="inline"><mml:msup><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:math></inline-formula>), <bold>(e)</bold> height (H1: 0–10, H2: 10–15, H3:
<inline-formula><mml:math id="M13" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> 15 m), <bold>(f)</bold> building position (P1: corner, P2: extreme, P3:
internal, P4: isolated), <bold>(g)</bold> ground slope condition (GS1: ridge, GS2: plain,
GS3: moderate slope, GS4: steep slope), <bold>(h)</bold> regularity in plan and elevation
(IRe: irregular, Re: regular), <bold>(i)</bold> roof type (RT1: heavy, no thrust, RT2:
heavy thrust, RT3: light, no thrust, RT4: light thrust), <bold>(j)</bold> construction
material (CM1: poor-quality masonry, CM2: good-quality masonry, CM3:
poor-quality mixed frame masonry, CM4: good-quality mixed frame masonry,
CM5: reinforced concrete frame, CM6: reinforced concrete wall, CM7: steel
frames, CM8: other), and <bold>(k)</bold> macroseismic intensity.</p></caption>
        <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://nhess.copernicus.org/articles/23/3199/2023/nhess-23-3199-2023-f02.png"/>

      </fig>

      <p id="d1e1719">To consider spatially distributed ground motion, the original DaDO data are
supplemented with the main-event macroseismic intensities (MSIs) provided by
the United States Geological Survey (USGS) ShakeMap tool
(Wald et al., 2005). MSIs given in
terms of modified Mercalli intensities are considered and assigned to
buildings based on their location. The distribution of MSI values in the
database is shown in Fig. 2k.</p>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Method</title>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Machine learning models</title>
      <?pagebreak page3203?><p id="d1e1737">Ghimire et al. (2022) applied classification- and regression-based machine
learning models to the damage observed after the 2015 Gorkha
earthquake, Nepal (NPC, 2015). The main concepts for method selection, the
definition of the dataset for training and testing, and the representation
of model performance are presented here.<?xmltex \hack{\newpage}?></p>
      <p id="d1e1741">To develop the heuristic damage assessment model, the damage grades are
considered the target feature. The damage grades are discrete labels,
from DG0 to DG5. The three most advanced classification and regression machine
learning algorithms were selected: random forest (RFC) and random forest regression (RFR)
(Breiman, 2001), gradient boosting classification (GBC) and
gradient boosting regression (GBR) (Friedman, 1999), and extreme gradient boosting
classification (XGBC) and extreme gradient boosting regression (XGBR) (Chen and
Guestrin, 2016). A label (or class) was thus assigned to the categorical
response variables (DG) for the classification-based machine learning
models. For the regression-based machine learning models, DG is converted
into a continuous variable to minimize misclassifications (Ghimire et al.,
2022). For the regression-based machine learning models, DG is converted
into a continuous variable as tested by Ghimire et al. (2022): first, the
damage grades were ordered and considered a continuous variable ranging
between 0 (DG0) and 5 (DG5). Because the regression model outputs a real
value between 0  and 5 and not an integer, we rounded the output (real
number) to the nearest integer to plot the confusion matrix. However, the
error matrices were computed without rounding the model outputs to the
nearest integer.</p>
      <?pagebreak page3204?><p id="d1e1744">Building features and macroseismic intensities were considered input
features. A one-hot encoding technique was used to convert the categorical
features (i.e. ground slope condition, building position, roof type,
construction material) into binary values (1 or 0), resulting in 28 input
variables (Table 2). No input features were removed from the dataset: some
building features (e.g. number of storeys and height) may be correlated, but
we assumed that the presence of correlated features does not impact the
overall performance of these machine learning methods (Ghimire et al.,
2022). No specific data cleaning methods were applied to the DaDO database.</p>
      <p id="d1e1747">The machine learning algorithms from the scikit-learn package developed in
Python (Pedregosa et al., 2011) were applied. The machine learning models
were trained and tested on the randomly selected training (60 % of the
dataset) and testing (40 % of the dataset) subsets of data, considering a
single earthquake dataset or the whole DaDO dataset. The testing subset was
kept hidden from the model during the training phase.</p>
</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Machine learning model efficacy</title>
      <p id="d1e1758">The efficacy of the heuristic damage assessment model (i.e. its ability to
predict damage to a satisfactory or expected degree) was analysed in three
stages: comparison of the efficacy of the machine learning models using
metrics, analysis of specific issues related to machine learning using the
selected models, and application of the heuristic model to the whole DaDO
dataset.</p>
<sec id="Ch1.S3.SS2.SSS1">
  <label>3.2.1</label><title>First stage: model selection</title>
      <p id="d1e1768">In the first stage, only the L'Aquila 2009 portfolio was considered for the
training and testing phases. This is the largest dataset in terms of the
number of buildings and was obtained using the AeDES survey format
(Baggio et al., 2007;
Dolce et al., 2019). Model efficacy was provided by a confusion matrix,
which represents model prediction compared with the so-called “ground
truth” value. Accuracy was then represented on the confusion matrix by the
ratio of the number of correctly predicted DGs to the total number of
observed values per DG (<inline-formula><mml:math id="M14" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">DG</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>).</p>
      <p id="d1e1782">Total accuracy (<inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) was computed as the ratio of the number of
correctly predicted DGs to the total number of observed values. <inline-formula><mml:math id="M16" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and
<inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">DG</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> values close to 1 indicate high efficacy. Moreover, the
quantitative statistical error was also calculated as the mean of the
absolute value of errors (MAE) and the mean squared error (MSE) (MAE and MSE
values close to 0 indicate high efficacy). For classification-based machine
learning models, the ordinal value of the DG was used to calculate the MAE
and MSE scores directly. For the regression-based machine learning models,
the output DG values were rounded to the nearest integer for the accuracy
scores plotted for the confusion matrix but not for the MAE and MSE value
calculations.</p>
</sec>
<sec id="Ch1.S3.SS2.SSS2">
  <label>3.2.2</label><title>Second stage: machine-learning-related issues</title>
      <p id="d1e1826">In the second stage, the best heuristic model for damage assessment was
selected based on the highest efficacy and used to analyse and test
specific issues related to machine learning: (1) the imbalance distribution
of DGs in DaDO; (2) the performance of the selected model when only some
basic, but accurately assessed, building features are considered (i.e.
number of storeys, location, age, floor area); and (3) the simplification of
the heuristic model, in the sense that DGs are grouped into a
traffic-light-based classification (i.e. green, yellow, and red,
corresponding to DG0 <inline-formula><mml:math id="M18" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> DG1, DG2 <inline-formula><mml:math id="M19" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> DG3, and DG4 <inline-formula><mml:math id="M20" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> DG5, respectively). In the
second stage, the issues related to machine learning were first analysed
using the L'Aquila 2009 portfolio. The whole DaDO dataset was then used.</p>
</sec>
<sec id="Ch1.S3.SS2.SSS3">
  <label>3.2.3</label><title>Third stage: application to the whole DaDO portfolio and comparison
with RISK-UE</title>
      <p id="d1e1858">In the third stage, several learning and testing sequences were considered,
with the idea of moving to an operational configuration in which past
information is used to predict damage from future earthquakes: either
learning based on a portfolio of damage caused by one earthquake and tested
on another portfolio or learning based on a series of damage portfolios and
tested on the portfolio of damage caused by an earthquake placed in the
chronological continuity of the earthquake sequence considered. In this
stage, the efficacy of the heuristic damage assessment model was analysed by
comparing the prediction values with the so-called ground truth values
through the error distribution as follows:
              <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M21" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ε</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub><mml:mfenced open="(" close=")"><mml:mi mathvariant="italic">%</mml:mi></mml:mfenced><mml:mo>=</mml:mo><mml:mfenced open="(" close=")"><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow><mml:mi>N</mml:mi></mml:mfrac></mml:mstyle></mml:mfenced><mml:mo>⋅</mml:mo><mml:mn mathvariant="normal">100</mml:mn><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
            where <inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the total number of buildings at a given error level
(difference between observed and predicted DGs) and <inline-formula><mml:math id="M23" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> is the total number of
buildings in the damage portfolio.</p>
      <p id="d1e1912">In this stage, the efficacy of the heuristic damage assessment model was
compared with the conventional damage prediction framework proposed by the
RISK-UE method (Milutinovic and Trendafiloski, 2003). The RISK-UE
method assigns a vulnerability index (<inline-formula><mml:math id="M24" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi mathvariant="normal">V</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) to a building, based on its
construction material and structural properties (e.g. height, building age,
position, regularities, geographic location). For a given level of
seismic demand (MSI), the mean damage (<inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) and the probability (<inline-formula><mml:math id="M26" display="inline"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) of observing a given damage level <inline-formula><mml:math id="M27" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> (<inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo></mml:mrow></mml:math></inline-formula> 0 to 5) are given by

                  <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M29" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E2"><mml:mtd><mml:mtext>2</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2.5</mml:mn><mml:mfenced open="[" close="]"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>+</mml:mo><mml:mi>tan⁡</mml:mi><mml:mi>h</mml:mi><mml:mfenced close=")" open="("><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi mathvariant="normal">MSI</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">6.25</mml:mn><mml:msub><mml:mi mathvariant="normal">I</mml:mi><mml:mi mathvariant="normal">V</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mn mathvariant="normal">13.1</mml:mn></mml:mrow><mml:mn mathvariant="normal">2.3</mml:mn></mml:mfrac></mml:mstyle></mml:mfenced></mml:mrow></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E3"><mml:mtd><mml:mtext>3</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mn mathvariant="normal">5</mml:mn><mml:mi mathvariant="normal">!</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mi mathvariant="normal">!</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mn mathvariant="normal">5</mml:mn><mml:mo>-</mml:mo><mml:mi>k</mml:mi><mml:mi mathvariant="normal">!</mml:mi></mml:mrow></mml:mfenced></mml:mrow></mml:mfrac></mml:mstyle><mml:msup><mml:mfenced close=")" open="("><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub></mml:mrow><mml:mn mathvariant="normal">5</mml:mn></mml:mfrac></mml:mstyle></mml:mfenced><mml:mn mathvariant="normal">5</mml:mn></mml:msup><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub></mml:mrow><mml:mn mathvariant="normal">5</mml:mn></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced><mml:mrow><mml:mn mathvariant="normal">5</mml:mn><mml:mo>-</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msup><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

              Herein, comparing the heuristic model and the RISK-UE method amounts to
considering the following steps, based on the equations given by RISK-UE.
<list list-type="bullet"><list-item>
      <p id="d1e2099"><italic>Step 1.</italic> The buildings in the training and testing datasets are grouped into
different classes according to construction material.</p></list-item><list-item>
      <p id="d1e2105"><italic>Step 2.</italic> For a given building class in the training dataset, computation of the following is performed:
<list list-type="custom"><list-item><label>-</label>
      <p id="d1e2112"><italic>Step 2.1.</italic> The mean damage (<inline-formula><mml:math id="M30" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) using the observed damage distribution
at a given MSI value is given by<disp-formula id="Ch1.E4" content-type="numbered"><label>4</label><mml:math id="M31" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow><mml:mn mathvariant="normal">5</mml:mn></mml:msubsup><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mi>k</mml:mi><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p></list-item><list-item><label>-</label>
      <p id="d1e2162"><italic>Step 2.2.</italic> The vulnerability index (<inline-formula><mml:math id="M32" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi mathvariant="normal">V</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) with the <inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> obtained in step 2.1 is given
by<disp-formula id="Ch1.E5" content-type="numbered"><label>5</label><mml:math id="M34" display="block"><mml:mrow><?xmltex \hack{\hbox\bgroup\fontsize{8.8}{8.8}\selectfont$\displaystyle}?><mml:msub><mml:mi mathvariant="normal">I</mml:mi><mml:mi mathvariant="normal">V</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">6.25</mml:mn></mml:mfrac></mml:mstyle><mml:mfenced open="[" close="]"><mml:mrow><mml:mn mathvariant="normal">13.1</mml:mn><mml:mo>-</mml:mo><mml:mi mathvariant="normal">MSI</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">2.3</mml:mn><mml:mfenced close=")" open="("><mml:mrow><mml:mi>tan⁡</mml:mi><mml:msup><mml:mi>h</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub></mml:mrow><mml:mn mathvariant="normal">2.5</mml:mn></mml:mfrac></mml:mstyle><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced><?xmltex \hack{$\egroup}?><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p></list-item></list></p></list-item><list-item>
      <p id="d1e2252"><italic>Step 3.</italic> For the same building class in the test dataset, calculation of the following is performed:
<list list-type="custom"><list-item><label>-</label>
      <p id="d1e2259"><italic>Step 3.1.</italic> The mean damage (<inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) of Eq. (2) for a given MSI value with the
value of <inline-formula><mml:math id="M36" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi mathvariant="normal">V</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> obtained in step 2.2 is calculated.</p></list-item><list-item><label>-</label>
      <p id="d1e2287"><italic>Step 3.2.</italic> The damage probability (<inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) of Eq. (3) with the value of <inline-formula><mml:math id="M38" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>
obtained in step 3.1 is calculated.</p></list-item><list-item><label>-</label>
      <p id="d1e2315"><italic>Step 3.3.</italic> The distribution of buildings in each damage grade within a range of
MSI values observed in the test dataset is calculated as<disp-formula id="Ch1.E6" content-type="numbered"><label>6</label><mml:math id="M39" display="block"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">pred</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mo>∑</mml:mo><mml:mi mathvariant="normal">MSI</mml:mi></mml:msub><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi mathvariant="normal">obs</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">MSI</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula><mml:math id="M40" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi mathvariant="normal">obs</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">MSI</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is the total number of buildings observed in the
test set for a given MSI value.</p></list-item><list-item><label>-</label>
      <p id="d1e2377"><italic>Step 3.4.</italic> The absolute error (<inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ε</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) in each damage level <inline-formula><mml:math id="M42" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> is given
by<disp-formula id="Ch1.E7" content-type="numbered"><label>7</label><mml:math id="M43" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ε</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfenced open="|" close="|"><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">obs</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">k</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">pred</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mi>N</mml:mi></mml:mfrac></mml:mstyle></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula><mml:math id="M44" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">obs</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is the total number of buildings observed in the given
damage grade <inline-formula><mml:math id="M45" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>.</p></list-item></list></p></list-item></list>
Similarly, the heuristic damage assessment model was also compared with the
mean damage relationship (Eq. 4) applied to the test set. Thus, for each
building class in the test set, the error value (Eq. 7) for each DG was
computed from the <inline-formula><mml:math id="M46" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> of the observed damage using Eq. (4), the
probability <inline-formula><mml:math id="M47" display="inline"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> of obtaining a given DG <inline-formula><mml:math id="M48" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> (<inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo></mml:mrow></mml:math></inline-formula> 0 to 5) using Eq. (3),
and the distribution of buildings in each DG <inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi mathvariant="normal">pred</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> for a given MSI
value using Eq. (6).</p>
</sec>
</sec>
</sec>
<?pagebreak page3205?><sec id="Ch1.S4">
  <label>4</label><title>Result</title>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>First stage: model selection</title>
      <p id="d1e2540">The efficacy of the regression (RFR, GBR, XGBR) and classification (RFC,
GBC, XGBC) machine learning models trained and tested on the randomly
selected 60 % (training set) and 40 % (test set) of the 2009-L'Aquila
earthquake building damage portfolio is summarized in Table 3. The
hyperparameters indicated in Table 3 were chosen after tests performed by
Ghimire et al. (2022). The regression-based machine learning models RFR, GBR,
and XGBR yielded similar MSE scores (1.22, 1.22, and 1.21) and accuracy
scores (<inline-formula><mml:math id="M51" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub><mml:mo>=</mml:mo></mml:mrow></mml:math></inline-formula> 0.49, 0.50, and 0.50), considering the five DGs of the
EMS-98 scale. In the confusion matrix (Fig. 3a: RFR, Fig. 3b: GBR, Fig. 3c: XGBR), the accuracy <inline-formula><mml:math id="M52" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">DG</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> values show that the efficacy of these
models is higher for the lower DGs (around 60 % for DG0 and 55 % for
DG1) and lower for the higher DGs (6 % and 1 % of the buildings are
correctly classified in DG4 and DG5, respectively).</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T3" specific-use="star"><?xmltex \currentcnt{3}?><label>Table 3</label><caption><p id="d1e2570">Summary of optimized hyperparameter parameters, accuracy <inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>,
and quantitative statistical error values for the regression-based and
classification-based machine learning methods in the test set. The
parameters are the hyperparameters chosen for the machine learning models
(the other model parameters not mentioned here are the default parameters in
the scikit-learn documentation; Pedregosa et
al., 2011). The best accuracy and error values are indicated in bold. The
optimum hyperparameters were selected thanks to <inline-formula><mml:math id="M54" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-fold cross-validation
(10-fold), by randomly selecting a percentage for training and percentage for testing,
for different combinations of hyperparameters and the optimum evaluated in
terms of performance metrics on testing is finally selected.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="5">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Method</oasis:entry>
         <oasis:entry colname="col2">Parameters</oasis:entry>
         <oasis:entry colname="col3">Accuracy</oasis:entry>
         <oasis:entry colname="col4">MSE</oasis:entry>
         <oasis:entry colname="col5">MAE</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M55" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">RFR</oasis:entry>
         <oasis:entry colname="col2">n_estimators <inline-formula><mml:math id="M56" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 1000</oasis:entry>
         <oasis:entry colname="col3">0.49</oasis:entry>
         <oasis:entry colname="col4">1.22</oasis:entry>
         <oasis:entry colname="col5">0.77</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">max_depth <inline-formula><mml:math id="M57" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 25</oasis:entry>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">GBR</oasis:entry>
         <oasis:entry colname="col2">n_estimators <inline-formula><mml:math id="M58" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 1000</oasis:entry>
         <oasis:entry colname="col3">0.50</oasis:entry>
         <oasis:entry colname="col4">1.22</oasis:entry>
         <oasis:entry colname="col5">0.77</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">max_depth <inline-formula><mml:math id="M59" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 10</oasis:entry>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">learning_ rate <inline-formula><mml:math id="M60" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.01</oasis:entry>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">XGBR</oasis:entry>
         <oasis:entry colname="col2">n_estimators <inline-formula><mml:math id="M61" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 1000</oasis:entry>
         <oasis:entry colname="col3">0.50</oasis:entry>
         <oasis:entry colname="col4"><bold>1.21</bold></oasis:entry>
         <oasis:entry colname="col5">0.76</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">max_depth <inline-formula><mml:math id="M62" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 10</oasis:entry>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">learning_ rate <inline-formula><mml:math id="M63" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.01</oasis:entry>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RFC</oasis:entry>
         <oasis:entry colname="col2">n_estimators <inline-formula><mml:math id="M64" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 1000</oasis:entry>
         <oasis:entry colname="col3">0.57</oasis:entry>
         <oasis:entry colname="col4">1.86</oasis:entry>
         <oasis:entry colname="col5">0.77</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">max_depth <inline-formula><mml:math id="M65" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 25</oasis:entry>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">GBC</oasis:entry>
         <oasis:entry colname="col2">n_estimators <inline-formula><mml:math id="M66" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 1000</oasis:entry>
         <oasis:entry colname="col3">0.58</oasis:entry>
         <oasis:entry colname="col4">1.80</oasis:entry>
         <oasis:entry colname="col5">0.77</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">max_depth <inline-formula><mml:math id="M67" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 10</oasis:entry>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">learning_ rate <inline-formula><mml:math id="M68" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.01</oasis:entry>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">XGBC</oasis:entry>
         <oasis:entry colname="col2">n_estimators <inline-formula><mml:math id="M69" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 1000</oasis:entry>
         <oasis:entry colname="col3"><bold>0.59</bold></oasis:entry>
         <oasis:entry colname="col4">1.78</oasis:entry>
         <oasis:entry colname="col5"><bold>0.74</bold></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">max_depth <inline-formula><mml:math id="M70" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 10</oasis:entry>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">learning_ rate <inline-formula><mml:math id="M71" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.01</oasis:entry>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table><?xmltex \gdef\@currentlabel{3}?></table-wrap>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3" specific-use="star"><?xmltex \currentcnt{3}?><?xmltex \def\figurename{Figure}?><label>Figure 3</label><caption><p id="d1e3021">Normalized confusion matrix between predicted and observed DGs.
The values given in each main diagonal cell are the accuracy scores
<inline-formula><mml:math id="M72" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">DG</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. All values are also represented by the colour scale.</p></caption>
          <?xmltex \igopts{width=369.885827pt}?><graphic xlink:href="https://nhess.copernicus.org/articles/23/3199/2023/nhess-23-3199-2023-f03.png"/>

        </fig>

      <p id="d1e3042">For the classification-based machine learning models, the XGBC model ([MSE,
<inline-formula><mml:math id="M73" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>] <inline-formula><mml:math id="M74" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> [1.78, 0.59]) was more effective than the RFC ([MSE,
<inline-formula><mml:math id="M75" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>]<inline-formula><mml:math id="M76" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> [1.86, 0.57]) and GBC ([MSE, <inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>] <inline-formula><mml:math id="M78" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> [1.80, 0.58])
models, considering the EMS-98 scale. In the confusion matrix (Fig. 3d: RFC,
Fig. 3e: GBC, Fig. 3f: XGBC), the accuracy <inline-formula><mml:math id="M79" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">DG</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> values also show
higher model efficacy for the lower DGs (86 % for DG0 and 39 % for DG1)
and lower efficacy for the higher DGs (5 %, 23 %, 12 %, and 17 %
buildings correctly classified in DG2, DG3, DG4, and DG5, respectively).</p>
      <?pagebreak page3206?><p id="d1e3111">The classification-based machine learning models thus yielded slightly
better predictive efficacy, but it was still lower than in recent studies using
other datasets (Ghimire et al., 2022; Harirchian et al., 2021; Mangalathu et
al., 2020; Roeslin et al., 2020; Stojadinović et al., 2021). The high
classification error in the higher DGs could be related to the
characteristics of the building portfolio and the imbalance of DG
distribution. Among the classification methods, the XGBC model showed
slightly higher classification efficacy; the XGBC model was therefore
selected for the next stages, stages 2 and 3.<?xmltex \hack{\newpage}?></p>
</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>Second stage: issues related to machine learning</title>
<sec id="Ch1.S4.SS2.SSS1">
  <label>4.2.1</label><title>Imbalance distribution of the DGs in DaDO</title>
      <p id="d1e3130">The efficacy of the heuristic damage assessment model depends on the
distribution of target features in the training dataset. This can lead to
low prediction efficacy, especially for minority classes
(Estabrooks and Japkowicz, 2001;
Japkowicz and Stephen, 2002; Branco et al., 2017; Ghimire et al., 2022). The
previous section reports significant misclassification associated with the
highest DGs for all classification- and regression-based models (Fig. 3),
i.e. for the DGs with the lowest number of buildings (Fig. 2a). The
efficacy of the XGBC model is analysed below, addressing the class-imbalance
issue with data resampling techniques applied to the training phase and
considering the L'Aquila 2009 portfolio.</p>
      <p id="d1e3133">Four strategies to solve the class-imbalance issue were tested:
<list list-type="custom"><list-item><label>a.</label>
      <p id="d1e3138">random undersampling – randomly selecting the number of data entries in
each class equal to the number of data entries in the minority class (DG4 in
our case);</p></list-item><list-item><label>b.</label>
      <p id="d1e3142">random oversampling – randomly replacing the number of data entries in
each class equal to the number of data entries in the majority class (DG0 in
our case);</p></list-item><list-item><label>c.</label>
      <p id="d1e3146">the synthetic minority oversampling technique (SMOTE) – creating an equal
number of data entries in each class by generating synthetic samples by
interpolating the neighbouring data in the minority class;</p></list-item><list-item><label>d.</label>
      <p id="d1e3150">a combination of oversampling and undersampling methods – oversampling of
the minority class using the SMOTE method, followed by the edited nearest
neighbours (ENN) undersampling method to eliminate data that are
misclassified by their three nearest neighbours (SMOTE-ENN).</p></list-item></list></p>
      <p id="d1e3153">Figure 4 shows the confusion matrices of the four strategies considered for
the class-imbalance issue. Compared with Fig. 3f (i.e. XGBC), the effects
of addressing the issue of imbalance were as follows:
<list list-type="custom"><list-item><label>a.</label>
      <p id="d1e3158"><italic>Undersampling (Fig. 4a).</italic> The <inline-formula><mml:math id="M80" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">DG</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> value increased by 20 %/22 %/26 % for
DG2/DG4/DG5 and decreased by 29 % for DG0.</p></list-item><list-item><label>b.</label>
      <p id="d1e3175"><italic>Oversampling (Fig. 4b).</italic> The <inline-formula><mml:math id="M81" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">DG</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> value increased by 11 %/16 %/18 % for
DG2/DG4/DG5 and decreased by 13 % for DG0.</p></list-item><list-item><label>c.</label>
      <p id="d1e3192"><italic>SMOTE (Fig. 4c).</italic> The <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">DG</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> value increased by 4 %/1 %/4 % for DG2/DG4/DG5
and decreased by 3 % for DG0.</p></list-item><list-item><label>d.</label>
      <p id="d1e3209"><italic>SMOTE-ENN (Fig. 4d).</italic> The <inline-formula><mml:math id="M83" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">DG</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> value increased by 13 %/9 %/8 % for
DG2/DG4/DG5 and decreased by 25 % for DG0.</p></list-item></list></p>
      <p id="d1e3225">The <inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, MAE, and MSE scores are given in Table 4 with the associated
effects.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T4" specific-use="star"><?xmltex \currentcnt{4}?><label>Table 4</label><caption><p id="d1e3243">Scores of the accuracy <inline-formula><mml:math id="M85" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, MSE, and MAE metrics in the test
set considering the imbalance issue and their variation <inline-formula><mml:math id="M86" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula> compared
with values without consideration of the imbalance.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="7">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right" colsep="1"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right" colsep="1"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Method</oasis:entry>
         <oasis:entry rowsep="1" namest="col2" nameend="col3" align="center" colsep="1">Accuracy <inline-formula><mml:math id="M87" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry rowsep="1" namest="col4" nameend="col5" align="center" colsep="1">MSE </oasis:entry>
         <oasis:entry rowsep="1" namest="col6" nameend="col7" align="center">MAE </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">Scores</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M88" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Score</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M89" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">Score</oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M90" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula></oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Undersampling</oasis:entry>
         <oasis:entry colname="col2">0.26</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M91" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.33</oasis:entry>
         <oasis:entry colname="col4">1.24</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M92" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.34</oasis:entry>
         <oasis:entry colname="col6">1.20</oasis:entry>
         <oasis:entry colname="col7">0.46</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Oversampling</oasis:entry>
         <oasis:entry colname="col2">0.53</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M93" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.06</oasis:entry>
         <oasis:entry colname="col4">2.13</oasis:entry>
         <oasis:entry colname="col5">0.35</oasis:entry>
         <oasis:entry colname="col6">0.86</oasis:entry>
         <oasis:entry colname="col7">0.12</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SMOTE</oasis:entry>
         <oasis:entry colname="col2">0.57</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M94" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.02</oasis:entry>
         <oasis:entry colname="col4">1.87</oasis:entry>
         <oasis:entry colname="col5">0.09</oasis:entry>
         <oasis:entry colname="col6">0.77</oasis:entry>
         <oasis:entry colname="col7">0.03</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SMOTE-ENN</oasis:entry>
         <oasis:entry colname="col2">0.49</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M95" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.10</oasis:entry>
         <oasis:entry colname="col4">2.28</oasis:entry>
         <oasis:entry colname="col5">0.50</oasis:entry>
         <oasis:entry colname="col6">0.93</oasis:entry>
         <oasis:entry colname="col7">0.19</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table><?xmltex \gdef\@currentlabel{4}?></table-wrap>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4" specific-use="star"><?xmltex \currentcnt{4}?><?xmltex \def\figurename{Figure}?><label>Figure 4</label><caption><p id="d1e3486">Confusion matrices for the four methods to solve the DG imbalance
issue in DaDO. The values given in each main diagonal cell are the
accuracy scores <inline-formula><mml:math id="M96" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">DG</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. All values are also represented by the colour
scale.</p></caption>
            <?xmltex \igopts{width=369.885827pt}?><graphic xlink:href="https://nhess.copernicus.org/articles/23/3199/2023/nhess-23-3199-2023-f04.png"/>

          </fig>

      <p id="d1e3506">In conclusion, the random oversampling method improves prediction in the
minority class without significantly decreasing prediction in the majority
class. The random oversampling method was therefore applied in this study.</p>
</sec>
<?pagebreak page3208?><sec id="Ch1.S4.SS2.SSS2">
  <label>4.2.2</label><title>Testing the XGBC model with basic features</title>
      <p id="d1e3517">This section begins by exploring the importance of each feature in the
heuristic damage assessment model applied to the L'Aquila 2009 portfolio. We
used the Shapley additive explanations (SHAP) method developed by Lundberg
and Lee (2017). The SHAP method compares the
efficacy of the model with and without considering each input feature to
measure its average impact, provided in terms of mean absolute SHAP values.<?xmltex \hack{\newpage}?></p>
      <p id="d1e3521">Figure 5a shows the average SHAP value associated with each feature
considered in this study as a function of DG. The most weighted features are
building age, location (latitude and longitude), material (poor-quality
masonry, reinforced concrete (RC) frame), MSI, roof type, floor area, and height. Interestingly,
the mean SHAP values are dependent on the DG; i.e. the weight of the
feature is not linear depending on the DG considered – this is never taken
into account in vulnerability methods. For example, Scala et al. (2022) and Del Gaudio et al. (2021) observed a decrease in<?pagebreak page3209?> the
vulnerability of structures as the construction year increases, without
distinguishing the DG considered, which is not the case herein. Note also
that the importance score associated with the location feature can
indirectly capture variations in local geological properties and the
spatially distributed vulnerability associated with the built-up area of the
L'Aquila 2009 portfolio (e.g. the distinction between the historic town and
more modern urban areas). Furthermore, the average SHAP value obtained for
poor-quality masonry buildings for DG3, DG4, and DG5 confirms the same high
vulnerability of this typology as in the EMS-98 scale (Grünthal, 1998),
regardless of DG.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5" specific-use="star"><?xmltex \currentcnt{5}?><?xmltex \def\figurename{Figure}?><label>Figure 5</label><caption><p id="d1e3526"><bold>(a)</bold> Graphic representation of the importance scores associated
with the different input features considered for the XGBC model. The
features (the same as in Fig. 2) considered in this study are on the <inline-formula><mml:math id="M97" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> axis,
and the <inline-formula><mml:math id="M98" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> axis is the mean SHAP score according to DG. <bold>(b)</bold> Confusion
matrices considering the basic-features setting. The values given in each
main diagonal cell are the accuracy scores <inline-formula><mml:math id="M99" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">DG</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. All values are also
represented by the colour scale.</p></caption>
            <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://nhess.copernicus.org/articles/23/3199/2023/nhess-23-3199-2023-f05.png"/>

          </fig>

      <p id="d1e3566">Some basic features of the building (e.g. location, age, floor area, number
of storeys, height) are observed with a high mean SHAP value (Fig. 5a).
Compared with others, these five basic features can easily be collected from
the field or provided by national census databases, for example. Figure 5b
shows the efficacy of the heuristic damage assessment model using XGBC
trained with a set of easily accessible building features (i.e.
basic-features setting: geographic location, floor area, number of stories,
height, age, MSI), after addressing the class-imbalance issue using the
random oversampling method. Compared with Fig. 4b (considering all features
and named as the full-features setting), the XGBC model with the
basic-features setting (Fig. 5b) gives almost the same efficacy, with only a
6 % average reduction in the accuracy scores.</p>
</sec>
<sec id="Ch1.S4.SS2.SSS3">
  <label>4.2.3</label><title>Testing the XGBC model with the traffic-light system for damage grades</title>
      <p id="d1e3577">In this section, a simplified version of the DG scale was used, in the sense
that the DGs are classified according to a traffic-light system (TLS) (i.e.
green G, yellow Y, and red R classes, corresponding to DG0 <inline-formula><mml:math id="M100" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> DG1, DG2 <inline-formula><mml:math id="M101" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> DG3,
and DG4 <inline-formula><mml:math id="M102" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> DG5, respectively), as monitored during post-earthquake emergency
situations
(Mangalathu
et al., 2020; Riedel et al., 2015; ATC, 2005; Bazzurro et al., 2004). For
the TLS-based damage classification, the XGBC model (after oversampling to
compensate for the imbalance issue) with the basic-features setting applied
to the L'Aquila 2009 portfolio (Fig. 6a) gives almost the same efficacy
compared to the full-features setting (Fig. 6b). For example, accuracy
values <inline-formula><mml:math id="M103" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">DG</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> using the basic-features setting and the
full-features setting were 0.76/0.34/0.56 and 0.82/0.36/0.54 for G/Y/R
classes, with the accuracy scores (<inline-formula><mml:math id="M104" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) of 0.68 and 0.72, respectively.
Mangalathu et al. (2020), Roeslin et al. (2020), and Harirchian et al. (2021) reported similar damage grade classification accuracy values of 0.66,
0.67, and 0.65, respectively.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6" specific-use="star"><?xmltex \currentcnt{6}?><?xmltex \def\figurename{Figure}?><label>Figure 6</label><caption><p id="d1e3625">Confusion matrices for <bold>(a)</bold> the basic-features setting and <bold>(b)</bold> the
full-features setting using the classification based on the traffic-light system (TLS),
grouping the EMS-98 damage grades (DGs) into three classes (green for no or
slight damage, yellow for moderate damage, and red for heavy damage). The
values given in each main diagonal cell are the accuracy scores <inline-formula><mml:math id="M105" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">DG</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.
All values are also represented by the colour scale.</p></caption>
            <?xmltex \igopts{width=369.885827pt}?><graphic xlink:href="https://nhess.copernicus.org/articles/23/3199/2023/nhess-23-3199-2023-f06.png"/>

          </fig>

      <p id="d1e3651">The efficacy of the heuristic damage assessment model using TLS-based damage
classification indicates that classifying damage into three classes is much
easier for the machine learning model compared with the six-class
classification system (EMS-98 damage classification). This is also observed
during damage surveys in the field, which sometimes find it hard to
distinguish between the intermediate damage grades, such as between DG2 and DG3 or between DG3 and
DG4. Similar observations have been reported in previous studies by
Guettiche
et al. (2017), Harirchian et al. (2021), Riedel et al. (2015), Roeslin et
al. (2020), and Stojadinović et al. (2021).</p>
</sec>
<sec id="Ch1.S4.SS2.SSS4">
  <label>4.2.4</label><title>Testing the XGBC model with the whole dataset</title>
      <p id="d1e3662">The efficacy of the XGBC model was tested using a dataset with six building
damage portfolios, excluding the 1980-Irpinia building damage portfolio. The
XGBC model was trained and tested on the randomly selected 60 % (training
set) and 40 % (test set) of the dataset for EMS-98/TLS damage
classification, with two sets of features (full-features setting and
basic-features setting), applying the random oversampling method to
compensate for class-imbalance issues. Figure 7 shows the associated confusion
matrix.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F7" specific-use="star"><?xmltex \currentcnt{7}?><?xmltex \def\figurename{Figure}?><label>Figure 7</label><caption><p id="d1e3667">Confusion matrices for EMS-98 <bold>(a, b)</bold> and TLS (green
for no or slight damage, yellow for moderate damage, red for heavy damage) <bold>(c, d)</bold> damage
classification systems using the full-features setting <bold>(a, c)</bold> and the basic-features setting <bold>(b, d)</bold>. The
values given in each main diagonal cell are the accuracy scores <inline-formula><mml:math id="M106" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">DG</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.
All values are also represented by the colour scale.</p></caption>
            <?xmltex \igopts{width=369.885827pt}?><graphic xlink:href="https://nhess.copernicus.org/articles/23/3199/2023/nhess-23-3199-2023-f07.png"/>

          </fig>

      <p id="d1e3699">The basic-features setting resulted in a similar level of damage prediction
compared with the full-features setting for both EMS-98-based and TLS-based damage
classification systems. For EMS-98 damage classification (Fig. 7a, b), the
accuracy <inline-formula><mml:math id="M107" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">DG</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> scores indicated in the confusion matrices are almost the
same for the basic-features setting and the full-features setting.
Furthermore, the accuracy <inline-formula><mml:math id="M108" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and MAE scores are also almost the same
(0.45 and 1.08 for the basic-features setting and 0.48 and 0.95 for the
full-features setting).</p>
      <p id="d1e3725">Likewise, for TLS-based damage classification (Fig. 7c, d), the accuracy
values <inline-formula><mml:math id="M109" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">DG</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> for the basic-features setting/full-features setting
are almost the same, with similar accuracy <inline-formula><mml:math id="M110" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">T</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and MAE scores (0.63/0.45
and 0.67/0.39, respectively).</p>
</sec>
</sec>
<sec id="Ch1.S4.SS3">
  <label>4.3</label><title>Third stage: application to the whole DaDO portfolio and comparison with
Risk-UE</title>
      <?pagebreak page3210?><p id="d1e3759">In this section, the efficacy of the heuristic damage assessment model was
considered for building damage predictions, without considering the time
frame of the earthquakes. Two scenarios were considered: (1) a single
building damage portfolio was used for training, and the model was then
tested on the others (named single-single), in situations using a single
portfolio to predict future damage, and (2) some building damage portfolios
were used for training but testing was performed on a single portfolio
(named aggregate-single); i.e. more damage portfolios were
used as a training set to predict the damage caused by the next earthquake.
The model XGBC was applied with the basic-features setting (number of
storeys, building age, floor area, height, MSI for EMS-98) and EMS-98-based and
TLS-based damage classification.<?xmltex \hack{\newpage}?></p>
<sec id="Ch1.S4.SS3.SSS1">
  <label>4.3.1</label><title>Single-single scenario</title>
      <p id="d1e3770">First, a series of building damage portfolios, concerning earthquakes
occurring in northern or southern Italy and of different magnitudes, were
used for training and testing:
<list list-type="custom"><list-item><label>i.</label>
      <p id="d1e3775">training set E3 and test sets E1, E5, and E7;</p></list-item><list-item><label>ii.</label>
      <p id="d1e3779">training set E5 and test sets E1, E3, and E7;</p></list-item><list-item><label>iii.</label>
      <p id="d1e3783">training set E7 and test sets E1, E3, and E5.</p></list-item></list>
Figure 8 shows the distribution of correct DG classification (i.e.
<inline-formula><mml:math id="M111" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="italic">ε</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> in percent given by Eq. 1) observed for each building
for the EMS-98 damage grade (Fig. 8a) and the TLS (Fig. 8b) systems. The <inline-formula><mml:math id="M112" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> axis
represents the incremental error in the damage grade (e.g. 1 corresponds to
the <inline-formula><mml:math id="M113" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula> of the damage grade between observation and prediction, regardless of
the DG considered).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F8" specific-use="star"><?xmltex \currentcnt{8}?><?xmltex \def\figurename{Figure}?><label>Figure 8</label><caption><p id="d1e3818">Distribution of the classification value (<inline-formula><mml:math id="M114" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="italic">ε</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>
in percent given by Eq. 1) for <bold>(a)</bold> EMS-98-based and <bold>(b)</bold> TLS-based damage
classification using XGBC machine learning models and considering a single
damage portfolio to predict a single portfolio (single-single scenario). The
colour bar indicates the associated value in each cell. The <inline-formula><mml:math id="M115" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> values are the
difference between the DG observed and the DG predicted, regardless of the
DG considered.</p></caption>
            <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://nhess.copernicus.org/articles/23/3199/2023/nhess-23-3199-2023-f08.png"/>

          </fig>

      <p id="d1e3855">For the EMS-98 damage scale, correct classification (<inline-formula><mml:math id="M116" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> value centred on 0)
in the range of 31 % to 48 % was found, depending on the training/test
datasets. The error distribution is quite wide with incorrect predictions
of <inline-formula><mml:math id="M117" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula>1 DG in the range of <inline-formula><mml:math id="M118" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula>13 % to 35 %. Remarkably, when considering
the E1 portfolio (Irpinia 1980), for which the post-earthquake inventory was
based on another form, as the test set, the error is larger. The predictions
at <inline-formula><mml:math id="M119" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula>1 DG (i.e. the sum of the <inline-formula><mml:math id="M120" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> values in Fig. 8a between <inline-formula><mml:math id="M121" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>1 and <inline-formula><mml:math id="M122" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>1)
were 70.5 %, 69.9 %, and 72.8 % with portfolios E3, E5, and E7 as the
test set, respectively, for an average of 71 %. For the other portfolios,
the average of the predictions at <inline-formula><mml:math id="M123" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula>1 DG was 77 %, 78 %, and 77 %,
respectively, for portfolios E5, E3, and E7 as the test set. This tendency
was also observed for the TLS damage system (Fig. 8b). In this case, the
classification of the E1 portfolio was correct on average (average of
<inline-formula><mml:math id="M124" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> values centred on 0) at 63 % and equal to 72 %, 73 %, and 70.5 %
for the test on portfolios E5, E3, and E7. For both damage scales, the
distributions were skewed, with a larger number of predictions being
underestimated (positive <inline-formula><mml:math id="M125" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> values), which is certainly a consequence of the choice
of machine learning models, their implementation (including imbalance
issues), the distribution of input and target features considered, or all of these aspects.
The interest of the machine learning model is also to have a relevant
representation of the errors and limits of these methods.</p>
</sec>
<?pagebreak page3211?><sec id="Ch1.S4.SS3.SSS2">
  <label>4.3.2</label><title>Aggregate-single scenario</title>
      <p id="d1e3937">Secondly, several aggregated building damage portfolio scenarios were
considered to predict a single earthquake, thus testing whether the
prediction was improved by increasing the number of post-earthquake damage
observations. Three scenarios were tested. They are represented in Fig. 9,
applying the EMS-98 damage grade (Fig. 9a) and the TLS (Fig. 9b):
<list list-type="custom"><list-item><label>i.</label>
      <p id="d1e3942">training set E2 <inline-formula><mml:math id="M126" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> E3 <inline-formula><mml:math id="M127" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> E4 <inline-formula><mml:math id="M128" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> E6 (shown as E2346) and test sets E1, E5, and E7;</p></list-item><list-item><label>ii.</label>
      <p id="d1e3967">training set E2 <inline-formula><mml:math id="M129" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> E4 <inline-formula><mml:math id="M130" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> E5 <inline-formula><mml:math id="M131" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> E6 (shown as E2456) and test sets E1, E3, and E7;</p></list-item><list-item><label>iii.</label>
      <p id="d1e3992">training set E2 <inline-formula><mml:math id="M132" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> E4 <inline-formula><mml:math id="M133" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> E6 <inline-formula><mml:math id="M134" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> E7 (shown as E2467) and test sets E1, E3, and E5.</p></list-item></list>
For the EMS-98 damage scale, correct classification (<inline-formula><mml:math id="M135" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> value centred on 0)
in the range of 27 % to 49 % was found, depending on the training/test
datasets. As in Fig. 8, using the E1 (Irpinia 1980) earthquake for testing
scored lower regardless of the portfolio used for training (28.7 %,
27.2 %, and 27.4 % prediction accuracy). With E1 as the test set, the
predictions at <inline-formula><mml:math id="M136" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula>1 DG (i.e. the sum of the <inline-formula><mml:math id="M137" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> values in Fig. 9a between
<inline-formula><mml:math id="M138" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>1 and <inline-formula><mml:math id="M139" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula>1) were 65.7 %, 63.8 %, and 62.4 % considering the E2346,
E2456, and E2467 portfolios as the training set, respectively, for an average
of 64 % (compared with the 70 % score for the single portfolio scenario,
Fig. 8a). Other scenarios were also tested by aggregating the building
damage portfolios differently (not presented herein), leading to two
main conclusions: (1) the quality and homogeneity of the input data (i.e.
building features) affect the efficacy of the heuristic model and (2) this
efficacy is limited and not improved by increasing the number of building
damage observations, with a score (excluding E1) of between 40 % and 49 %
(<inline-formula><mml:math id="M140" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> value centred on 0) and up to 78 % (average of the two scenarios, Figs. 8a and 9a) at <inline-formula><mml:math id="M141" display="inline"><mml:mo>±</mml:mo></mml:math></inline-formula>1 DG. Considering the TLS damage scale (Fig. 9b), a
damage prediction efficacy of about 72 % was obtained (compared with
72 % in Fig. 8b), but no significant improvement was observed when
the number of damaged buildings in the training portfolio was increased. For
EMS-98 and TLS, the distributions were skewed, with a larger number of
predictions being underestimated (positive <inline-formula><mml:math id="M142" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> values).</p>

      <?xmltex \floatpos{p}?><fig id="Ch1.F9" specific-use="star"><?xmltex \currentcnt{9}?><?xmltex \def\figurename{Figure}?><label>Figure 9</label><caption><p id="d1e4076">Distribution of the classification value (<inline-formula><mml:math id="M143" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="italic">ε</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>
in percent given by Eq. 1) for <bold>(a)</bold> EMS-98-based and <bold>(b)</bold> TLS-based damage
classification using XGBC machine learning models and considering an
aggregate damage portfolio to predict a single portfolio (aggregate-single
scenario). The colour bar indicates the associated value in each cell. The
<inline-formula><mml:math id="M144" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> values are the difference between the DG observed and the DG predicted,
regardless of the DG considered.</p></caption>
            <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://nhess.copernicus.org/articles/23/3199/2023/nhess-23-3199-2023-f09.png"/>

          </fig>

      <p id="d1e4113">In conclusion, the heuristic damage assessment model based on the
XGBC model gives a better score for TLS damage assessment than for the
EMS-98 damage scale. The TLS system also allows for quick assessment of
damage on a large scale such as a city or region from an operational point
of view.</p>
</sec>
<?pagebreak page3212?><sec id="Ch1.S4.SS3.SSS3">
  <label>4.3.3</label><title>Comparing efficacy with the RISK-UE model</title>
      <p id="d1e4124">The efficacy of the heuristic damage assessment model was then compared with
conventional damage prediction methods, i.e. RISK-UE and the mean damage
relationship (Eqs. 2 to 7), considering the basic-features settings. For
RISK-UE, mean damage <inline-formula><mml:math id="M145" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> (Eq. 4) was computed using the training set
and the vulnerability index <inline-formula><mml:math id="M146" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi mathvariant="normal">V</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> for each building (Eq. 5). A vulnerability
index was then attributed to all the buildings<?pagebreak page3213?> in each class defined
according to building features. The vulnerability indexes were then
attributed to every building in the test set; mean damage (<inline-formula><mml:math id="M147" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) was
computed with Eq. (2) and then DG distribution with Eq. (3), before being
compared with the damage portfolio used for testing. Finally, the
distribution of the mean damage observed (Eq. 4) was compared with the
distribution of damage directly on the test set, using Eq. (3).</p>
      <p id="d1e4160">Figure 10 shows the distribution of absolute errors associated with the
RISK-UE, mean damage relationship, and XGBC methods (with and without
compensation for the class-imbalance issue) trained on earthquake building
damage portfolio E5 and tested on E3. For EMS-98 damage classification (Fig. 10a), the XGBC model (without compensation for class-imbalance issues)
resulted in a level of absolute errors similar to that of the RISK-UE and/or
mean damage relationship, except for DG0 (24 %). Random oversampling to
compensate for the class-imbalance issues improved the distribution of
errors for the XGBC model (errors less than 8 %, except for DG1 at 13 %).</p>

      <?xmltex \floatpos{p}?><fig id="Ch1.F10" specific-use="star"><?xmltex \currentcnt{10}?><?xmltex \def\figurename{Figure}?><label>Figure 10</label><caption><p id="d1e4165">Comparison of the efficacy of the heuristic model with the
conventional model considering the DaDO portfolio (training set: E5; test
set: E3) for <bold>(a)</bold> EMS-98-based and <bold>(b)</bold> TLS-based damage classification. The <inline-formula><mml:math id="M148" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> axis
is the damage grade, and the <inline-formula><mml:math id="M149" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> axis is the percentage of absolute error
(<inline-formula><mml:math id="M150" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ε</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> in percent given by Eq. 7). The blue bar corresponds to the
mean damage relationship; the red bar corresponds to the RISK-UE method; and the
green and orange bars correspond to the heuristic model without (XGBC<inline-formula><mml:math id="M151" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:math></inline-formula>)
and with (XGBC<inline-formula><mml:math id="M152" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>) compensation for the class-imbalance issues,
respectively.</p></caption>
            <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://nhess.copernicus.org/articles/23/3199/2023/nhess-23-3199-2023-f10.png"/>

          </fig>

      <p id="d1e4225">For TLS-based damage classification, the XGBC model also resulted in a
similar level of errors compared with the mean damage relationship and/or
RISK-UE methods (Fig. 10b), except for the green class (no or slight damage,
17.04 %). Compensation for class-imbalance issues slightly improved the
distribution of errors for the XGBC model with a 2 % drop in errors for
the green (no/slight damage) and yellow (moderate damage) classes.</p>
      <p id="d1e4228">Figure 11 shows the distribution of absolute errors trained using the E2456
portfolio and tested on the E3 portfolio. For EMS-98 damage classification
(Fig. 11a), the XGBC model (without compensation for class-imbalance issues)
resulted in a level of errors similar to that of the RISK-UE and/or mean
damage relationship; errors were highest for DG0 with 15.15 %. With
compensation for the class-imbalance issues, the XGBC model achieved a
slightly lower error distribution for DG0 (5 %) and DG3 (4 %); however,
for other damage grades, the error value increased significantly (DG1:
11 %, DG2: 12 %, DG4: 7 %, DG5: 2 %). For TLS-based damage
classification, the distribution of absolute errors was similar for both the
XGBC model and the mean damage relationship and/or RISK-UE methods (Fig. 11b). The highest absolute error value was associated with the green (no or
slight damage) class of buildings (16.40 %). Compensation for the
class-imbalance issues slightly increased the error distribution for the
XGBC model, with nearly 5 % for buildings in the green (no or slight damage) and
red (heavy damage) classes.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F11" specific-use="star"><?xmltex \currentcnt{11}?><?xmltex \def\figurename{Figure}?><label>Figure 11</label><caption><p id="d1e4233">Comparison of the efficacy of the heuristic model with the
conventional model considering the DaDO portfolio (training set: E2456; test
set: E3) for <bold>(a)</bold> EMS-98-based and <bold>(b)</bold> TLS-based damage classification. The <inline-formula><mml:math id="M153" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> axis
is the damage grade, and the <inline-formula><mml:math id="M154" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> axis is the percentage of absolute error
(<inline-formula><mml:math id="M155" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ε</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> in percent given by Eq. 7). The blue bar corresponds to the
mean damage relationship; the red bar corresponds to the RISK-UE method; and the
green and orange bars correspond to the heuristic model without (XGBC<inline-formula><mml:math id="M156" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:math></inline-formula>)
and with (XGBC<inline-formula><mml:math id="M157" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>) compensation for the class-imbalance issues,
respectively.</p></caption>
            <?xmltex \igopts{width=426.791339pt}?><graphic xlink:href="https://nhess.copernicus.org/articles/23/3199/2023/nhess-23-3199-2023-f11.png"/>

          </fig>

      <p id="d1e4292">These results show that the heuristic building damage model based on the
XGBC model, trained using building damage portfolios with the
basic-features setting, provides a reasonable estimation of potential
damage, particularly with TLS-based damage classification.</p>
</sec>
</sec>
</sec>
<sec id="Ch1.S5">
  <label>5</label><title>Discussion</title>
      <p id="d1e4306">Previous studies have aimed to test a machine learning framework for seismic
building damage assessment (e.g. Mangalathu et al., 2020; Roeslin et al.,
2020; Harirchian et al., 2021; Ghimire et al., 2022). They evaluated various
machine learning and data balancing methods to classify earthquake damage to
buildings. However, these studies (Mangalathu et al., 2020; Roeslin et al.,
2020; Harirchian et al., 2021) had limitations such as limited data samples, limited
damage classes, and building characteristics limited to a spatial coverage
and range of seismic demand values. Ghimire et al. (2022) also used a larger
building damage database but did not investigate the importance of input
features as a function of damage levels and did not compare machine learning
with conventional damage assessment methods.</p>
      <p id="d1e4309">This study aims to go beyond previous studies by testing advanced machine
learning methods and data resampling techniques using the unique DaDO
dataset collected from several major earthquakes in Italy. This database
covers a wide range of seismic damage and seismic demands of a specific
region, including undamaged buildings. Most importantly, this study
highlights the importance of input features according to the degrees of
damage and finally compares the machine learning models with a classical
damage prediction model (RISK-UE). The machine learning models achieved
comparable accuracy to the RISK-UE method. In addition, TLS-based damage
classification, using red for heavily damaged, yellow for moderate damage,
and green for no to slight damage, could be appropriate when the information
about undamaged buildings is unavailable during model training.</p>
      <p id="d1e4312">Indeed, it is worth noting that the importance of the input features used in
the learning process changes with the degree of damage: this indicates that
each feature may have a contribution to the damage that changes with the
damage level. Thus, the weight of each feature does not depend linearly on
the degree of damage, which is not considered in conventional vulnerability
methods.</p>
      <?pagebreak page3215?><p id="d1e4315">The prediction of seismic damage by machine learning remains until now has been
tested on geographically limited data. The damage distribution is strongly
influenced by region-specific factors such as construction quality and
regional typologies, implementation of seismic regulations, and hazard level.
Therefore, machine-learning-based models can only work well in regions with
comparable characteristics, and a host-to-target transfer of these models
should be studied. In addition, the distribution of damage is often
imbalanced, impacting the performance of machine learning models by
assigning higher weights to the features of the majority class. However,
data balancing methods like random oversampling can reduce bias caused by
imbalanced data during the training phase, but they may also introduce
overfitting issues depending on the distribution of input and target
features. Thus, integrating data from a wider range of input features and
earthquake damage from different regions, relying on a host-to-target
strategy, could help achieve a more natural balance of datasets and lead to
less biased results. Moreover, the machine learning methods only train on
the data available in the learning phase that reflect the building
portfolio in the study area. The importance of the features contributing to
the damage could thus be modulated and would require a host-to-target
adjustment for the application of the model to another urban zone/seismic
region.</p>
      <p id="d1e4319">However, the machine learning models trained and tested on the DaDO dataset
resulted in similar damage prediction accuracy values to those reported in existing
literature using different models and datasets with different combinations
of input features. This might suggest that the uncertainty related to
building vulnerability in damage classification may be smaller than the
primary source of uncertainty related to the hazard component (such as
ground motion, fault rupture, or slip duration).</p>
      <p id="d1e4322">In recent years, there has been a proliferation of open building data, such
as the OpenStreetMap-based dynamic global exposure model (Schorlemmer et
al., 2020) and building damage datasets after an earthquake (such as DaDO).
We must therefore continue this paradigm shift initiated by Riedel et al. (2014, 2015), which consisted in identifying the exposure data available and with
as much certainty as possible and in finding the most effective relationships for
estimating the damage, unlike conventional approaches, which proposed
established and robust methods but relied on data that were not available or were difficult to collect. The global dynamic exposure model will make
it possible to meet the challenge of modelling exposure on a larger scale with
available data, using a tool capable of integrating this large volume of
data. Machine learning methods are one such rapidly growing tool that can
aid in exposure classification and damage prediction by leveraging readily
available information. It is therefore necessary to continue in this
direction in order to evaluate the performance of the methods and their pros
and cons for maximum efficacy of the prediction of damage.</p>
      <p id="d1e4325">Future works will therefore have to address several key issues that have
been discussed here but that need to be further investigated. For example,
the weight of the input features varies according to the level of damage,
but one can question the systematization of this observation whatever the
dataset and feature considered. The efficiency of the selected models
and the management of imbalance data remain to be explored, in particular by
verifying regional independence. Taking advantage of the increasing
abundance of exposure data and post-seismic observations, the imbalanced
feature distribution and observed damage levels could be solved by
aggregating datasets independent of the exposure and hazard contexts of the
regions, once the host-to-target transfer of the models has been resolved.
Finally, key input features (still not yet identified) describing hazard or
vulnerability may be unexplored, and incorporating them into the models may
improve the accuracy of damage classification.</p>
</sec>
<sec id="Ch1.S6" sec-type="conclusions">
  <label>6</label><title>Conclusions</title>
      <p id="d1e4336">In this study, we explored the efficacy of machine learning models trained
using DaDO post-earthquake building damage portfolios. We compared six
machine learning models: RFC, GBC, XGBC, RFR, GBR, and XGBR. These models
were trained on a number of building features (location, number of storeys,
age, floor area, height, position, construction material, regularity, roof
type, ground slope condition) and ground motion intensity defined in terms
of macroseismic<?pagebreak page3216?> intensity. The classification models performed slightly
better than the regression methods, and the XGBC model was ultimately found
to be the most efficient model for this dataset. To solve the imbalance
issue concerning observed damage, the random oversampling method was applied
to the training dataset to improve the efficacy of the heuristic damage
assessment model by rectifying the skewed distribution of the target
features (DGs).</p>
      <p id="d1e4339">Surprisingly, we found that the weight of the most important building
feature evolves according to DG; i.e. the weight of the feature for damage
prediction changes depending on the DG considered. This is not taken into
account in conventional methods.</p>
      <p id="d1e4342">The basic-features setting (i.e. considering the number of storeys, age, floor
area, height, and macroseismic intensity, which are accurately evaluated for
the existing building portfolio) gave the same accuracy (0.68) as the
full-features setting (0.72) with the TLS-based damage classification
method. For training and testing, the homogeneity of the information in the
portfolios is a key issue for the definition of a highly effective machine
learning model, as shown by the data from the E1 earthquake (Irpinia-1990).
However, the efficacy of the model reaches a limit which is not improved by
increasing the number of damaged buildings in the portfolio used as the
training set, for example. For damage prediction, this type of heuristic
model results in approximately 75 % correct classification. Other authors
(e.g. Riedel et al., 2014, 2015; Ghimire
et al., 2022) have already reached this same conclusion by increasing the
percentage of the training set compared with the test set.</p>
      <p id="d1e4345">Despite this limit threshold, the level of accuracy achieved remains similar
to that attained by conventional methods, such as RISK-UE and the mean
damage relationship, for the basic-features settings and TLS-based damage
classification (error values less than 17 %). Machine learning models
trained on post-earthquake building damage portfolios could provide a
reasonable estimation of damage for a different region with similar building
portfolios, after host-to-target adjustment.</p>
      <p id="d1e4349">Some variability may have been introduced into the damage prediction model
due to the framework defined to translate the original damage scale to the
EMS-98 damage scale and because, in the DaDO database, the year of
construction and the floor area of each building are provided as interval
values and missing locations of buildings have been replaced with the location
of local administrative centres. The latter can lead to a smoothing of the
macroseismic intensities to be considered for each structure and also
affect the distance to the earthquake. Similarly, the building damage
surveys were carried out after the seismic sequence, which includes
aftershocks as well as the mainshock, whereas the MSI input corresponds to
the mainshock from the USGS ShakeMap. All these issues may reduce the
efficacy of the heuristic model and its limit threshold. Addressing these
issues could improve the damage prediction performance of machine learning
models.</p>
</sec>

      
      </body>
    <back><notes notes-type="codeavailability"><title>Code availability</title>

      <p id="d1e4356">The machine learning models were developed using scikit-learn documentation,
and the value of hyperparameters used are provided in Table 3 (<uri>https://scikit-learn.org/stable/install.html</uri>, Pedregosa et al., 2011).</p>
  </notes><notes notes-type="dataavailability"><title>Data availability</title>

      <p id="d1e4365">The data used in this study are available in the Database of Observed Damage
(DaDO) web-based GIS platform of the Italian Civil Protection Department, developed by the
Eucentre Foundation (<uri>https://egeos.eucentre.it/danno_osservato/web/danno_osservato?lang=EN</uri>, Dolce et al., 2019).</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d1e4374">SG: conceptualization, methodology, data preparation,
investigation, visualization, and draft preparation. PG:
conceptualization, investigation, visualization, supervision, and review and
editing draft. AP: conceptualization, supervision, and review and editing
draft. DS: conceptualization, supervision, and review and
editing draft.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e4380">The contact author has declared that none of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d1e4386">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.</p>
  </notes><notes notes-type="sistatement"><title>Special issue statement</title>

      <p id="d1e4392">This article is part of the special issue “Harmonized seismic hazard and risk assessment for Europe”. It is not associated with a conference.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e4398">The authors thank the editor and the reviewers. Adrien Pothon and Philippe Guéguen thank the AXA Research Fund supporting the project New Probabilistic Assessment of Seismic Hazard, Losses and Risks in Strong Seismic Prone Regions. Philippe Guéguen thanks LabEx OSUG@2020 (Investissements d’avenir – ANR10LABX56).</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d1e4404">This research has been supported by the URBASIS-EU project (H2020-MSCA-ITN-2018, grant no. 813137), the Agence Nationale de la Recherche (grant no. ANR10LABX56), and the
AXA Research Fund (New Probabilistic Assessment of Seismic Hazard, Losses and Risks in Strong Seismic Prone Regions).</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e4410">This paper was edited by Helen Crowley and reviewed by Zoran Stojadinovic, Marta Faravelli, and one anonymous referee.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bib1"><label>1</label><?label 1?><mixed-citation>
ATC: ATC-20-1, Field Manual: Postearthquake Safety Evaluation of Buildings
Second Edition, Applied Technology Council, Redwood City, California, ISBN  ATC20-1,
2005.</mixed-citation></ref>
      <ref id="bib1.bib2"><label>2</label><?label 1?><mixed-citation>Azimi, M., Eslamlou, A. D., and Pekcan, G.: Data-driven structural health
monitoring and damage detection through deep learning: State-of-the-art
review, Sensors, 20, 2778, <ext-link xlink:href="https://doi.org/10.3390/s20102778" ext-link-type="DOI">10.3390/s20102778</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib3"><label>3</label><?label 1?><mixed-citation>
Baggio, C., Bernardini, A., Colozza, R., Pinto, A. V., and Taucer, F.: Field
Manual for post-earthquake damage and safety assessment and short term
countermeasures (AeDES) Translation from Italian: Maria ROTA and Agostino
GORETTI, European Commission – Joint Research Centre – Institute for the Protection and Security of the Citizen, EUR 22868, 2007.</mixed-citation></ref>
      <ref id="bib1.bib4"><label>4</label><?label 1?><mixed-citation>Bazzurro, P., Cornell, C. A., Menun, C., and Motahari, M.: Guidelines for
seismic assessment of damaged buildings, in: 13th World Conference on
Earthquake Engineering, Vancouver, B.C., Canada, 74–76,
<ext-link xlink:href="https://doi.org/10.5459/bnzsee.38.1.41-49" ext-link-type="DOI">10.5459/bnzsee.38.1.41-49</ext-link>, 2004.</mixed-citation></ref>
      <ref id="bib1.bib5"><label>5</label><?label 1?><mixed-citation>
Branco, P., Ribeiro, R. P., Torgo, L., Krawczyk, B., and Moniz, N.: SMOGN: a
Pre-processing Approach for Imbalanced Regression, Proc. Mach. Learn. Res.,
74, 36–50, 2017.</mixed-citation></ref>
      <ref id="bib1.bib6"><label>6</label><?label 1?><mixed-citation>
Breiman, L.: Random Forests, Mach. Learn., 45, 5–32, 2001.</mixed-citation></ref>
      <ref id="bib1.bib7"><label>7</label><?label 1?><mixed-citation>Chen, T. and Guestrin, C.: XGBoost: A Scalable Tree Boosting System, in:
22nd acm sigkdd international conference on knowledge discovery and data
mining, San Francisco, CA, USA, 13–17 August 2016, 785–794, <ext-link xlink:href="https://doi.org/10.1145/2939672.2939785" ext-link-type="DOI">10.1145/2939672.2939785</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bib8"><label>8</label><?label 1?><mixed-citation>Daniell, J. E., Schaefer, A. M., Wenzel, F., and Tsang, H. H.: The global
role of earthquake fatalities in decision-making: earthquakes versus other
causes of fatalities, Proc. Sixt. World Conference Earthq. Eng. Santiago, Chile, 9–13 January 2017, <uri>http://www.wcee.nicee.org/wcee/article/16WCEE/WCEE2017-170.pdf</uri> (last access: 29 September 2023),
2017.</mixed-citation></ref>
      <ref id="bib1.bib9"><label>9</label><?label 1?><mixed-citation>Dolce, M., Speranza, E., Giordano, F., Borzi, B., Bocchi, F., Conte, C.,
Di Meo, A., Faravelli, M., and Pascale, V.: Observed damage database of past
italian earthquakes: The da.D.O. WebGIS, Bulletin of Geophyiscs and Oceanography,
60, 141–164, <ext-link xlink:href="https://doi.org/10.4430/bgta0254" ext-link-type="DOI">10.4430/bgta0254</ext-link>, 2019 (data available at: <uri>https://egeos.eucentre.it/danno_osservato/web/danno_osservato?lang=EN</uri>).</mixed-citation></ref>
      <ref id="bib1.bib10"><label>10</label><?label 1?><mixed-citation>Estabrooks, A. and Japkowicz, N.: A mixture-of-experts framework for
learning from imbalanced data sets, Lect. Notes Comput. Sc., 2189,
34–43, <ext-link xlink:href="https://doi.org/10.1007/3-540-44816-0_4" ext-link-type="DOI">10.1007/3-540-44816-0_4</ext-link>, 2001.</mixed-citation></ref>
      <ref id="bib1.bib11"><label>11</label><?label 1?><mixed-citation>FEMA: Hazus – MH 2.1 Multi-hazard Loss Estimation Methodology Earthquake,
<uri>https://www.fema.gov/sites/default/files/2020-09/fema_hazus_earthquake-model_technical-manual_2.1.pdf</uri> (last access: 29 September 2023), 2003.</mixed-citation></ref>
      <ref id="bib1.bib12"><label>12</label><?label 1?><mixed-citation>
Friedman, J. H. Greedy function approximation: a gradient boosting machine, Ann. Stat., 29, 1189–1232, 2001.</mixed-citation></ref>
      <ref id="bib1.bib13"><label>13</label><?label 1?><mixed-citation>
Del Gaudio, C., Scala, S. A., Ricci, P., and Verderame, G. M.: Evolution of
the seismic vulnerability of masonry buildings based on the damage data from
L'Aquila 2009 event, B. Earthq. Eng., 19, 4435–4470, 2021.</mixed-citation></ref>
      <ref id="bib1.bib14"><label>14</label><?label 1?><mixed-citation>Ghimire, S., Guéguen, P., Giffard-Roisin, S., and Schorlemmer, D.:
Testing machine learning models for seismic damage prediction at a regional
scale using building-damage dataset compiled after the 2015 Gorkha Nepal
earthquake, Earthq. Spectra, 38, 2970–2993, <ext-link xlink:href="https://doi.org/10.1177/87552930221106495" ext-link-type="DOI">10.1177/87552930221106495</ext-link>,
2022.</mixed-citation></ref>
      <ref id="bib1.bib15"><label>15</label><?label 1?><mixed-citation>Grünthal, G.: Escala Macro Sísmica Europea EMS – 98, 101 pp.,  <uri>https://www.franceseisme.fr/EMS98_Original_english.pdf</uri> (last access: 29 September 2023), 1998.</mixed-citation></ref>
      <ref id="bib1.bib16"><label>16</label><?label 1?><mixed-citation>Guéguen, P., Michel, C., and Lecorre, L.: A simplified approach for
vulnerability assessment in moderate-to-low seismic hazard regions:
Application to Grenoble (France), B. Earthq. Eng., 5, 467–490,
<ext-link xlink:href="https://doi.org/10.1007/s10518-007-9036-3" ext-link-type="DOI">10.1007/s10518-007-9036-3</ext-link>, 2007.</mixed-citation></ref>
      <ref id="bib1.bib17"><label>17</label><?label 1?><mixed-citation>Guettiche, A., Guéguen, P., and Mimoune, M.: Seismic vulnerability
assessment using association rule learning: application to the city of
Constantine, Algeria, Nat. Hazards, 86, 1223–1245,
<ext-link xlink:href="https://doi.org/10.1007/s11069-016-2739-5" ext-link-type="DOI">10.1007/s11069-016-2739-5</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib18"><label>18</label><?label 1?><mixed-citation>Harirchian, E., Kumari, V., Jadhav, K., Rasulzade, S., Lahmer, T., and Das,
R. R.: A synthesized study based on machine learning approaches for rapid
classifying earthquake damage grades to rc buildings, Appl. Sci., 11, 7540,
<ext-link xlink:href="https://doi.org/10.3390/app11167540" ext-link-type="DOI">10.3390/app11167540</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib19"><label>19</label><?label 1?><mixed-citation>Hegde, J. and Rokseth, B.: Applications of machine learning methods for
engineering risk assessment – A review, Safety Sci., 122, 104492,
<ext-link xlink:href="https://doi.org/10.1016/j.ssci.2019.09.015" ext-link-type="DOI">10.1016/j.ssci.2019.09.015</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib20"><label>20</label><?label 1?><mixed-citation>
Japkowicz, N. and Stephen, S.: The class imbalance problem A systematic
study fulltext.pdf, Intelligent Data Analysis, 6, 429–449, 2002.</mixed-citation></ref>
      <ref id="bib1.bib21"><label>21</label><?label 1?><mixed-citation>Kim, T., Song, J., and Kwon, O. S.: Pre- and post-earthquake regional loss
assessment using deep learning, Earthq. Eng. Struct. D., 49, 657–678,
<ext-link xlink:href="https://doi.org/10.1002/eqe.3258" ext-link-type="DOI">10.1002/eqe.3258</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib22"><label>22</label><?label 1?><mixed-citation>Lagomarsino, S. and Giovinazzi, S.: Macroseismic and mechanical models for
the vulnerability and damage assessment of current buildings, B. Earthq.
Eng., 4, 415–443, <ext-link xlink:href="https://doi.org/10.1007/s10518-006-9024-z" ext-link-type="DOI">10.1007/s10518-006-9024-z</ext-link>, 2006.</mixed-citation></ref>
      <ref id="bib1.bib23"><label>23</label><?label 1?><mixed-citation>Lagomarsino, S., Cattari, S., and Ottonelli, D.: The heuristic vulnerability
model: fragility curves for masonry buildings, Springer Netherlands,
3129–3163, <ext-link xlink:href="https://doi.org/10.1007/s10518-021-01063-7" ext-link-type="DOI">10.1007/s10518-021-01063-7</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib24"><label>24</label><?label 1?><mixed-citation>
Lundberg, S. M. and Lee, S.-I.: A Unified Approach to Interpreting Model
Predictions, in: 31st Conference on Neural Information Processing Systems 30, Annual Conference on Neural Information Processing Systems, 4–9 December 2017, Long Beach, CA, USA, 2017.</mixed-citation></ref>
      <ref id="bib1.bib25"><label>25</label><?label 1?><mixed-citation>Mangalathu, S. and Jeon, J.-S.: Regional Seismic Risk Assessment of
Infrastructure Systems through Machine Learning: Active Learning Approach,
J. Struct. Eng., 146, 04020269,
<ext-link xlink:href="https://doi.org/10.1061/(asce)st.1943-541x.0002831" ext-link-type="DOI">10.1061/(asce)st.1943-541x.0002831</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib26"><label>26</label><?label 1?><mixed-citation>Mangalathu, S., Sun, H., Nweke, C. C., Yi, Z., and Burton, H. V.:
Classifying earthquake damage to buildings using machine learning, Earthq.
Spectra, 36, 183–208, <ext-link xlink:href="https://doi.org/10.1177/8755293019878137" ext-link-type="DOI">10.1177/8755293019878137</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib27"><label>27</label><?label 1?><mixed-citation>Milutinovic, Z. and Trendafiloski, G.: Risk-UE An advanced approach to
earthquake risk scenarios with applications to different european towns,
Rep. to WP4 vulnerability Curr. Build., 1–83, <ext-link xlink:href="https://doi.org/10.1007/978-1-4020-3608-8_23" ext-link-type="DOI">10.1007/978-1-4020-3608-8_23</ext-link>, 2003.</mixed-citation></ref>
      <ref id="bib1.bib28"><label>28</label><?label 1?><mixed-citation>Ministerio de Vivienda y Urbanismo (MINVU): National Housing Reconstruction Program, <uri>https://www.preventionweb.net/files/28726_plandereconstruccinminvu.pdf</uri> (last access: 29 September 2023), 2010 (in Spanish).</mixed-citation></ref>
      <ref id="bib1.bib29"><label>29</label><?label 1?><mixed-citation>Morfidis, K. and Kostinakis, K.: Approaches to the rapid seismic damage
prediction of r/c buildings using artificial neural networks, Eng. Struct.,
165, 120–141, <ext-link xlink:href="https://doi.org/10.1016/j.engstruct.2018.03.028" ext-link-type="DOI">10.1016/j.engstruct.2018.03.028</ext-link>, 2018.</mixed-citation></ref>
      <?pagebreak page3218?><ref id="bib1.bib30"><label>30</label><?label 1?><mixed-citation>Mouroux, P. and Le Brun, B.: Presentation of RISK-UE Project, B. Earthq.
Eng., 44,  323–339, <ext-link xlink:href="https://doi.org/10.1007/S10518-006-9020-3" ext-link-type="DOI">10.1007/S10518-006-9020-3</ext-link>, 2006.</mixed-citation></ref>
      <ref id="bib1.bib31"><label>31</label><?label 1?><mixed-citation>Ministere des Travaux Publics, Transports et Communications (MTPTC): Evaluation des
Bâtiments:
<uri>https://www.mtptc.gouv.ht/accueil/recherche/article_7.html</uri>, last access: 26 September 2023.</mixed-citation></ref>
      <ref id="bib1.bib32"><label>32</label><?label 1?><mixed-citation>NPC: Post disaster needs assessment, <uri>https://www.npc.gov.np/images/category/PDNA_volume_BfinalVersion.pdf</uri> (last access: 27 September 2023), 2015.</mixed-citation></ref>
      <ref id="bib1.bib33"><label>33</label><?label 1?><mixed-citation>Pedregosa, F., Varoquaux, G., Buitinck, L., Louppe, G., Grisel, O., and
Mueller, A.: Scikit-learn, GetMobile Mob. Comput. Commun., 19, 29–33,
<ext-link xlink:href="https://doi.org/10.1145/2786984.2786995" ext-link-type="DOI">10.1145/2786984.2786995</ext-link>, 2011 (code available at: <uri>https://scikit-learn.org/stable/install.html</uri>).</mixed-citation></ref>
      <ref id="bib1.bib34"><label>34</label><?label 1?><mixed-citation>Riedel, I., Guéguen, P., Dunand, F., and Cottaz, S.: Macroscale
vulnerability assessment of cities using association rule learning, Seismol.
Res. Lett., 85, 295–305, <ext-link xlink:href="https://doi.org/10.1785/0220130148" ext-link-type="DOI">10.1785/0220130148</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bib35"><label>35</label><?label 1?><mixed-citation>Riedel, I., Guéguen, P., Dalla Mura, M., Pathier, E., Leduc, T., and
Chanussot, J.: Seismic vulnerability assessment of urban environments in
moderate-to-low seismic hazard regions using association rule learning and
support vector machine methods, Nat. Hazards, 76, 1111–1141,
<ext-link xlink:href="https://doi.org/10.1007/s11069-014-1538-0" ext-link-type="DOI">10.1007/s11069-014-1538-0</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bib36"><label>36</label><?label 1?><mixed-citation>Roeslin, S., Ma, Q., Juárez-Garcia, H., Gómez-Bernal, A., Wicker,
J., and Wotherspoon, L.: A machine learning damage prediction model for the
2017 Puebla-Morelos, Mexico, earthquake, Earthq. Spectra, 36, 314–339,
<ext-link xlink:href="https://doi.org/10.1177/8755293020936714" ext-link-type="DOI">10.1177/8755293020936714</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib37"><label>37</label><?label 1?><mixed-citation>Salehi, H. and Burgueño, R.: Emerging artificial intelligence methods in
structural engineering, Eng. Struct., 171, 170–189,
<ext-link xlink:href="https://doi.org/10.1016/j.engstruct.2018.05.084" ext-link-type="DOI">10.1016/j.engstruct.2018.05.084</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib38"><label>38</label><?label 1?><mixed-citation>Scala, S. A., Del Gaudio, C., and Verderame, G. M.: Influence of
construction age on seismic vulnerability of masonry buildings damaged after
2009 L'Aquila earthquake, Soil Dyn. Earthq. Eng., 157, 107199,
<ext-link xlink:href="https://doi.org/10.1016/J.SOILDYN.2022.107199" ext-link-type="DOI">10.1016/J.SOILDYN.2022.107199</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bib39"><label>39</label><?label 1?><mixed-citation>Schorlemmer, D., Beutin, T., Cotton, F., Garcia Ospina, N., Hirata, N., Ma, K.-F., Nievas, C., Prehn, K., and Wyss, M.: Global Dynamic Exposure and the OpenBuildingMap - A Big-Data and Crowd-Sourcing Approach to Exposure Modeling, EGU General Assembly 2020, Online, 4–8 May 2020, EGU2020-18920, https://doi.org/10.5194/egusphere-egu2020-18920, 2020.
 </mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bib40"><label>40</label><?label 1?><mixed-citation>Seo, J., Dueñas-Osorio, L., Craig, J. I., and Goodno, B. J.:
Metamodel-based regional vulnerability estimate of irregular steel
moment-frame structures subjected to earthquake events, Eng. Struct., 45,
585–597, <ext-link xlink:href="https://doi.org/10.1016/j.engstruct.2012.07.003" ext-link-type="DOI">10.1016/j.engstruct.2012.07.003</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bib41"><label>41</label><?label 1?><mixed-citation>Silva, V., Pagani, M., Schneider, J., and Henshaw, P.: Assessing Seismic
Hazard and Risk Globally for an Earthquake Resilient World, Contrib. Pap. to
GAR 2019, 24 pp., <uri>https://api.semanticscholar.org/CorpusID:208785925</uri> (last access: 29 September 2023), 2019.</mixed-citation></ref>
      <ref id="bib1.bib42"><label>42</label><?label 1?><mixed-citation>Silva, V., Brzev, S., Scawthorn, C., Yepes, C., Dabbeek, J., and Crowley,
H.: A Building Classification System for Multi-hazard Risk Assessment, Int.
J. Disast. Risk Sc., 13, 161–177,
<ext-link xlink:href="https://doi.org/10.1007/s13753-022-00400-x" ext-link-type="DOI">10.1007/s13753-022-00400-x</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bib43"><label>43</label><?label 1?><mixed-citation>Stojadinović, Z., Kovačević, M., Marinković, D., and
Stojadinović, B.: Rapid earthquake loss assessment based on machine
learning and representative sampling, Earthq. Spectra, 38, 152–177,
<ext-link xlink:href="https://doi.org/10.1177/87552930211042393" ext-link-type="DOI">10.1177/87552930211042393</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib44"><label>44</label><?label 1?><mixed-citation>Sun, H., Burton, H. V., and Huang, H.: Machine learning applications for
building structural design and performance assessment: State-of-the-art
review, J. Build. Eng., 33, 101816,
<ext-link xlink:href="https://doi.org/10.1016/j.jobe.2020.101816" ext-link-type="DOI">10.1016/j.jobe.2020.101816</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib45"><label>45</label><?label 1?><mixed-citation>Wald, D. J., Worden, B. C., Quitoriano, V., and Pankow, K. L.: ShakeMap
manual: technical manual, user's guide, and software guide, Techniques and
Methods, 134 pp., <ext-link xlink:href="https://doi.org/10.3133/tm12A1" ext-link-type="DOI">10.3133/tm12A1</ext-link>, 2005.</mixed-citation></ref>
      <ref id="bib1.bib46"><label>46</label><?label 1?><mixed-citation>Wang, C., Yu, Q., Law, K. H., McKenna, F., Yu, S. X., Taciroglu, E.,
Zsarnóczay, A., Elhaddad, W., and Cetiner, B.: Machine learning-based
regional scale intelligent modeling of building information for natural
hazard risk management, Autom. Constr., 122, 103474,
<ext-link xlink:href="https://doi.org/10.1016/j.autcon.2020.103474" ext-link-type="DOI">10.1016/j.autcon.2020.103474</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib47"><label>47</label><?label 1?><mixed-citation>Xie, Y., Ebad Sichani, M., Padgett, J. E., and DesRoches, R.: The promise of
implementing machine learning in earthquake engineering: A state-of-the-art
review, Earthq. Spectra, 36, 1769–1801,
<ext-link xlink:href="https://doi.org/10.1177/8755293020919419" ext-link-type="DOI">10.1177/8755293020919419</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib48"><label>48</label><?label 1?><mixed-citation>Xu, Y., Lu, X., Cetiner, B., and Taciroglu, E.: Real-time regional seismic
damage assessment framework based on long short-term memory neural network,
Comput. Civ. Infrastruct. Eng., 1–18, <ext-link xlink:href="https://doi.org/10.1111/mice.12628" ext-link-type="DOI">10.1111/mice.12628</ext-link>,
2020.</mixed-citation></ref>
      <ref id="bib1.bib49"><label>49</label><?label 1?><mixed-citation>Xu, Z., Wu Y., Qi, M., Zheng, M., Xiong, C., and Lu, X.: Prediction of
structural type for city-scale seismic damage simulation based on machine
learning, Appl. Sci., 10, 1795, <ext-link xlink:href="https://doi.org/10.3390/app10051795" ext-link-type="DOI">10.3390/app10051795</ext-link>, 2020.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Testing machine learning models for heuristic building damage assessment applied to the Italian Database of Observed Damage (DaDO)</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>1</label><mixed-citation>
      
ATC: ATC-20-1, Field Manual: Postearthquake Safety Evaluation of Buildings
Second Edition, Applied Technology Council, Redwood City, California, ISBN  ATC20-1,
2005.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>2</label><mixed-citation>
      
Azimi, M., Eslamlou, A. D., and Pekcan, G.: Data-driven structural health
monitoring and damage detection through deep learning: State-of-the-art
review, Sensors, 20, 2778, <a href="https://doi.org/10.3390/s20102778" target="_blank">https://doi.org/10.3390/s20102778</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>3</label><mixed-citation>
      
Baggio, C., Bernardini, A., Colozza, R., Pinto, A. V., and Taucer, F.: Field
Manual for post-earthquake damage and safety assessment and short term
countermeasures (AeDES) Translation from Italian: Maria ROTA and Agostino
GORETTI, European Commission – Joint Research Centre – Institute for the Protection and Security of the Citizen, EUR 22868, 2007.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>4</label><mixed-citation>
      
Bazzurro, P., Cornell, C. A., Menun, C., and Motahari, M.: Guidelines for
seismic assessment of damaged buildings, in: 13th World Conference on
Earthquake Engineering, Vancouver, B.C., Canada, 74–76,
<a href="https://doi.org/10.5459/bnzsee.38.1.41-49" target="_blank">https://doi.org/10.5459/bnzsee.38.1.41-49</a>, 2004.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>5</label><mixed-citation>
      
Branco, P., Ribeiro, R. P., Torgo, L., Krawczyk, B., and Moniz, N.: SMOGN: a
Pre-processing Approach for Imbalanced Regression, Proc. Mach. Learn. Res.,
74, 36–50, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>6</label><mixed-citation>
      
Breiman, L.: Random Forests, Mach. Learn., 45, 5–32, 2001.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>7</label><mixed-citation>
      
Chen, T. and Guestrin, C.: XGBoost: A Scalable Tree Boosting System, in:
22nd acm sigkdd international conference on knowledge discovery and data
mining, San Francisco, CA, USA, 13–17 August 2016, 785–794, <a href="https://doi.org/10.1145/2939672.2939785" target="_blank">https://doi.org/10.1145/2939672.2939785</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>8</label><mixed-citation>
      
Daniell, J. E., Schaefer, A. M., Wenzel, F., and Tsang, H. H.: The global
role of earthquake fatalities in decision-making: earthquakes versus other
causes of fatalities, Proc. Sixt. World Conference Earthq. Eng. Santiago, Chile, 9–13 January 2017, <a href="http://www.wcee.nicee.org/wcee/article/16WCEE/WCEE2017-170.pdf" target="_blank"/> (last access: 29 September 2023),
2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>9</label><mixed-citation>
      
Dolce, M., Speranza, E., Giordano, F., Borzi, B., Bocchi, F., Conte, C.,
Di Meo, A., Faravelli, M., and Pascale, V.: Observed damage database of past
italian earthquakes: The da.D.O. WebGIS, Bulletin of Geophyiscs and Oceanography,
60, 141–164, <a href="https://doi.org/10.4430/bgta0254" target="_blank">https://doi.org/10.4430/bgta0254</a>, 2019 (data available at: <a href="https://egeos.eucentre.it/danno_osservato/web/danno_osservato?lang=EN" target="_blank"/>).

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>10</label><mixed-citation>
      
Estabrooks, A. and Japkowicz, N.: A mixture-of-experts framework for
learning from imbalanced data sets, Lect. Notes Comput. Sc., 2189,
34–43, <a href="https://doi.org/10.1007/3-540-44816-0_4" target="_blank">https://doi.org/10.1007/3-540-44816-0_4</a>, 2001.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>11</label><mixed-citation>
      
FEMA: Hazus – MH 2.1 Multi-hazard Loss Estimation Methodology Earthquake,
<a href="https://www.fema.gov/sites/default/files/2020-09/fema_hazus_earthquake-model_technical-manual_2.1.pdf" target="_blank"/> (last access: 29 September 2023), 2003.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>12</label><mixed-citation>
      
Friedman, J. H. Greedy function approximation: a gradient boosting machine, Ann. Stat., 29, 1189–1232, 2001.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>13</label><mixed-citation>
      
Del Gaudio, C., Scala, S. A., Ricci, P., and Verderame, G. M.: Evolution of
the seismic vulnerability of masonry buildings based on the damage data from
L'Aquila 2009 event, B. Earthq. Eng., 19, 4435–4470, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>14</label><mixed-citation>
      
Ghimire, S., Guéguen, P., Giffard-Roisin, S., and Schorlemmer, D.:
Testing machine learning models for seismic damage prediction at a regional
scale using building-damage dataset compiled after the 2015 Gorkha Nepal
earthquake, Earthq. Spectra, 38, 2970–2993, <a href="https://doi.org/10.1177/87552930221106495" target="_blank">https://doi.org/10.1177/87552930221106495</a>,
2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>15</label><mixed-citation>
      
Grünthal, G.: Escala Macro Sísmica Europea EMS – 98, 101 pp.,  <a href="https://www.franceseisme.fr/EMS98_Original_english.pdf" target="_blank"/> (last access: 29 September 2023), 1998.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>16</label><mixed-citation>
      
Guéguen, P., Michel, C., and Lecorre, L.: A simplified approach for
vulnerability assessment in moderate-to-low seismic hazard regions:
Application to Grenoble (France), B. Earthq. Eng., 5, 467–490,
<a href="https://doi.org/10.1007/s10518-007-9036-3" target="_blank">https://doi.org/10.1007/s10518-007-9036-3</a>, 2007.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>17</label><mixed-citation>
      
Guettiche, A., Guéguen, P., and Mimoune, M.: Seismic vulnerability
assessment using association rule learning: application to the city of
Constantine, Algeria, Nat. Hazards, 86, 1223–1245,
<a href="https://doi.org/10.1007/s11069-016-2739-5" target="_blank">https://doi.org/10.1007/s11069-016-2739-5</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>18</label><mixed-citation>
      
Harirchian, E., Kumari, V., Jadhav, K., Rasulzade, S., Lahmer, T., and Das,
R. R.: A synthesized study based on machine learning approaches for rapid
classifying earthquake damage grades to rc buildings, Appl. Sci., 11, 7540,
<a href="https://doi.org/10.3390/app11167540" target="_blank">https://doi.org/10.3390/app11167540</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>19</label><mixed-citation>
      
Hegde, J. and Rokseth, B.: Applications of machine learning methods for
engineering risk assessment – A review, Safety Sci., 122, 104492,
<a href="https://doi.org/10.1016/j.ssci.2019.09.015" target="_blank">https://doi.org/10.1016/j.ssci.2019.09.015</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>20</label><mixed-citation>
      
Japkowicz, N. and Stephen, S.: The class imbalance problem A systematic
study fulltext.pdf, Intelligent Data Analysis, 6, 429–449, 2002.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>21</label><mixed-citation>
      
Kim, T., Song, J., and Kwon, O. S.: Pre- and post-earthquake regional loss
assessment using deep learning, Earthq. Eng. Struct. D., 49, 657–678,
<a href="https://doi.org/10.1002/eqe.3258" target="_blank">https://doi.org/10.1002/eqe.3258</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>22</label><mixed-citation>
      
Lagomarsino, S. and Giovinazzi, S.: Macroseismic and mechanical models for
the vulnerability and damage assessment of current buildings, B. Earthq.
Eng., 4, 415–443, <a href="https://doi.org/10.1007/s10518-006-9024-z" target="_blank">https://doi.org/10.1007/s10518-006-9024-z</a>, 2006.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>23</label><mixed-citation>
      
Lagomarsino, S., Cattari, S., and Ottonelli, D.: The heuristic vulnerability
model: fragility curves for masonry buildings, Springer Netherlands,
3129–3163, <a href="https://doi.org/10.1007/s10518-021-01063-7" target="_blank">https://doi.org/10.1007/s10518-021-01063-7</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>24</label><mixed-citation>
      
Lundberg, S. M. and Lee, S.-I.: A Unified Approach to Interpreting Model
Predictions, in: 31st Conference on Neural Information Processing Systems 30, Annual Conference on Neural Information Processing Systems, 4–9 December 2017, Long Beach, CA, USA, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>25</label><mixed-citation>
      
Mangalathu, S. and Jeon, J.-S.: Regional Seismic Risk Assessment of
Infrastructure Systems through Machine Learning: Active Learning Approach,
J. Struct. Eng., 146, 04020269,
<a href="https://doi.org/10.1061/(asce)st.1943-541x.0002831" target="_blank">https://doi.org/10.1061/(asce)st.1943-541x.0002831</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>26</label><mixed-citation>
      
Mangalathu, S., Sun, H., Nweke, C. C., Yi, Z., and Burton, H. V.:
Classifying earthquake damage to buildings using machine learning, Earthq.
Spectra, 36, 183–208, <a href="https://doi.org/10.1177/8755293019878137" target="_blank">https://doi.org/10.1177/8755293019878137</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>27</label><mixed-citation>
      
Milutinovic, Z. and Trendafiloski, G.: Risk-UE An advanced approach to
earthquake risk scenarios with applications to different european towns,
Rep. to WP4 vulnerability Curr. Build., 1–83, <a href="https://doi.org/10.1007/978-1-4020-3608-8_23" target="_blank">https://doi.org/10.1007/978-1-4020-3608-8_23</a>, 2003.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>28</label><mixed-citation>
      
Ministerio de Vivienda y Urbanismo (MINVU): National Housing Reconstruction Program, <a href="https://www.preventionweb.net/files/28726_plandereconstruccinminvu.pdf" target="_blank"/> (last access: 29 September 2023), 2010 (in Spanish).

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>29</label><mixed-citation>
      
Morfidis, K. and Kostinakis, K.: Approaches to the rapid seismic damage
prediction of r/c buildings using artificial neural networks, Eng. Struct.,
165, 120–141, <a href="https://doi.org/10.1016/j.engstruct.2018.03.028" target="_blank">https://doi.org/10.1016/j.engstruct.2018.03.028</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>30</label><mixed-citation>
      
Mouroux, P. and Le Brun, B.: Presentation of RISK-UE Project, B. Earthq.
Eng., 44,  323–339, <a href="https://doi.org/10.1007/S10518-006-9020-3" target="_blank">https://doi.org/10.1007/S10518-006-9020-3</a>, 2006.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>31</label><mixed-citation>
      
Ministere des Travaux Publics, Transports et Communications (MTPTC): Evaluation des
Bâtiments:
<a href="https://www.mtptc.gouv.ht/accueil/recherche/article_7.html" target="_blank"/>, last access: 26 September 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>32</label><mixed-citation>
      
NPC: Post disaster needs assessment, <a href="https://www.npc.gov.np/images/category/PDNA_volume_BfinalVersion.pdf" target="_blank"/> (last access: 27 September 2023), 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>33</label><mixed-citation>
      
Pedregosa, F., Varoquaux, G., Buitinck, L., Louppe, G., Grisel, O., and
Mueller, A.: Scikit-learn, GetMobile Mob. Comput. Commun., 19, 29–33,
<a href="https://doi.org/10.1145/2786984.2786995" target="_blank">https://doi.org/10.1145/2786984.2786995</a>, 2011 (code available at: <a href="https://scikit-learn.org/stable/install.html" target="_blank"/>).

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>34</label><mixed-citation>
      
Riedel, I., Guéguen, P., Dunand, F., and Cottaz, S.: Macroscale
vulnerability assessment of cities using association rule learning, Seismol.
Res. Lett., 85, 295–305, <a href="https://doi.org/10.1785/0220130148" target="_blank">https://doi.org/10.1785/0220130148</a>, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>35</label><mixed-citation>
      
Riedel, I., Guéguen, P., Dalla Mura, M., Pathier, E., Leduc, T., and
Chanussot, J.: Seismic vulnerability assessment of urban environments in
moderate-to-low seismic hazard regions using association rule learning and
support vector machine methods, Nat. Hazards, 76, 1111–1141,
<a href="https://doi.org/10.1007/s11069-014-1538-0" target="_blank">https://doi.org/10.1007/s11069-014-1538-0</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>36</label><mixed-citation>
      
Roeslin, S., Ma, Q., Juárez-Garcia, H., Gómez-Bernal, A., Wicker,
J., and Wotherspoon, L.: A machine learning damage prediction model for the
2017 Puebla-Morelos, Mexico, earthquake, Earthq. Spectra, 36, 314–339,
<a href="https://doi.org/10.1177/8755293020936714" target="_blank">https://doi.org/10.1177/8755293020936714</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>37</label><mixed-citation>
      
Salehi, H. and Burgueño, R.: Emerging artificial intelligence methods in
structural engineering, Eng. Struct., 171, 170–189,
<a href="https://doi.org/10.1016/j.engstruct.2018.05.084" target="_blank">https://doi.org/10.1016/j.engstruct.2018.05.084</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>38</label><mixed-citation>
      
Scala, S. A., Del Gaudio, C., and Verderame, G. M.: Influence of
construction age on seismic vulnerability of masonry buildings damaged after
2009 L'Aquila earthquake, Soil Dyn. Earthq. Eng., 157, 107199,
<a href="https://doi.org/10.1016/J.SOILDYN.2022.107199" target="_blank">https://doi.org/10.1016/J.SOILDYN.2022.107199</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>39</label><mixed-citation>
      
Schorlemmer, D., Beutin, T., Cotton, F., Garcia Ospina, N., Hirata, N., Ma, K.-F., Nievas, C., Prehn, K., and Wyss, M.: Global Dynamic Exposure and the OpenBuildingMap - A Big-Data and Crowd-Sourcing Approach to Exposure Modeling, EGU General Assembly 2020, Online, 4–8 May 2020, EGU2020-18920, https://doi.org/10.5194/egusphere-egu2020-18920, 2020.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>40</label><mixed-citation>
      
Seo, J., Dueñas-Osorio, L., Craig, J. I., and Goodno, B. J.:
Metamodel-based regional vulnerability estimate of irregular steel
moment-frame structures subjected to earthquake events, Eng. Struct., 45,
585–597, <a href="https://doi.org/10.1016/j.engstruct.2012.07.003" target="_blank">https://doi.org/10.1016/j.engstruct.2012.07.003</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>41</label><mixed-citation>
      
Silva, V., Pagani, M., Schneider, J., and Henshaw, P.: Assessing Seismic
Hazard and Risk Globally for an Earthquake Resilient World, Contrib. Pap. to
GAR 2019, 24 pp., <a href="https://api.semanticscholar.org/CorpusID:208785925" target="_blank"/> (last access: 29 September 2023), 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>42</label><mixed-citation>
      
Silva, V., Brzev, S., Scawthorn, C., Yepes, C., Dabbeek, J., and Crowley,
H.: A Building Classification System for Multi-hazard Risk Assessment, Int.
J. Disast. Risk Sc., 13, 161–177,
<a href="https://doi.org/10.1007/s13753-022-00400-x" target="_blank">https://doi.org/10.1007/s13753-022-00400-x</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>43</label><mixed-citation>
      
Stojadinović, Z., Kovačević, M., Marinković, D., and
Stojadinović, B.: Rapid earthquake loss assessment based on machine
learning and representative sampling, Earthq. Spectra, 38, 152–177,
<a href="https://doi.org/10.1177/87552930211042393" target="_blank">https://doi.org/10.1177/87552930211042393</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>44</label><mixed-citation>
      
Sun, H., Burton, H. V., and Huang, H.: Machine learning applications for
building structural design and performance assessment: State-of-the-art
review, J. Build. Eng., 33, 101816,
<a href="https://doi.org/10.1016/j.jobe.2020.101816" target="_blank">https://doi.org/10.1016/j.jobe.2020.101816</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>45</label><mixed-citation>
      
Wald, D. J., Worden, B. C., Quitoriano, V., and Pankow, K. L.: ShakeMap
manual: technical manual, user's guide, and software guide, Techniques and
Methods, 134 pp., <a href="https://doi.org/10.3133/tm12A1" target="_blank">https://doi.org/10.3133/tm12A1</a>, 2005.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>46</label><mixed-citation>
      
Wang, C., Yu, Q., Law, K. H., McKenna, F., Yu, S. X., Taciroglu, E.,
Zsarnóczay, A., Elhaddad, W., and Cetiner, B.: Machine learning-based
regional scale intelligent modeling of building information for natural
hazard risk management, Autom. Constr., 122, 103474,
<a href="https://doi.org/10.1016/j.autcon.2020.103474" target="_blank">https://doi.org/10.1016/j.autcon.2020.103474</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>47</label><mixed-citation>
      
Xie, Y., Ebad Sichani, M., Padgett, J. E., and DesRoches, R.: The promise of
implementing machine learning in earthquake engineering: A state-of-the-art
review, Earthq. Spectra, 36, 1769–1801,
<a href="https://doi.org/10.1177/8755293020919419" target="_blank">https://doi.org/10.1177/8755293020919419</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>48</label><mixed-citation>
      
Xu, Y., Lu, X., Cetiner, B., and Taciroglu, E.: Real-time regional seismic
damage assessment framework based on long short-term memory neural network,
Comput. Civ. Infrastruct. Eng., 1–18, <a href="https://doi.org/10.1111/mice.12628" target="_blank">https://doi.org/10.1111/mice.12628</a>,
2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>49</label><mixed-citation>
      
Xu, Z., Wu Y., Qi, M., Zheng, M., Xiong, C., and Lu, X.: Prediction of
structural type for city-scale seismic damage simulation based on machine
learning, Appl. Sci., 10, 1795, <a href="https://doi.org/10.3390/app10051795" target="_blank">https://doi.org/10.3390/app10051795</a>, 2020.

    </mixed-citation></ref-html>--></article>
