Abstract
The automated analysis of Terms and Conditions has gained attention in recent years, mainly due to its relevance to consumer protection. Well-structured data sets are the base for every analysis. While content extraction, in general, is a well-researched field and many open source libraries are available, our evaluation shows, that existing solutions cannot extract Terms and Conditions in sufficient quality, mainly because of their special structure. In this paper, we present an approach to extract the content and hierarchy of Terms and Conditions from German and English online shops. Our evaluation shows, that the approach outperforms the current state of the art. A python implementation of the approach is made available under an open license.
Original language | English |
---|---|
Title of host publication | Proceedings of The Fifth Workshop on e-Commerce and NLP (ECNLP 5) |
Editors | Shervin Malmasi, Oleg Rokhlenko, Nicola Ueffing, Ido Guy, Eugene Agichtein, Surya Kallumadi |
Place of Publication | Dublin, Ireland |
Publisher | Association for Computational Linguistics (ACL) |
Pages | 181-190 |
Number of pages | 10 |
ISBN (Electronic) | 978-1-955917-35-3 |
Publication status | Published - 1 May 2022 |
Event | The 5th Workshop on e-Commerce and NLP, ECNLP 2022 - Dublin, Ireland Duration: 26 May 2022 → 26 May 2022 |
Workshop
Workshop | The 5th Workshop on e-Commerce and NLP, ECNLP 2022 |
---|---|
Abbreviated title | ECNLP 2022 |
Country/Territory | Ireland |
City | Dublin |
Period | 26/05/22 → 26/05/22 |