FAQ-Gen: An automated system to generate domain-specific FAQs to aid content comprehension
Submitted: 2024-02-11
|Accepted: 2024-04-21
|Published: 2024-11-15
Copyright (c) 2024 Journal of Computer-Assisted Linguistic Research

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Downloads
Keywords:
Frequently Asked Questions, Natural Language Processing, Text-to-text Transformation, Transfer Learning, Natural Language Generation
Supporting agencies:
Stride.ai R&D Pvt Ltd, Bengaluru, India
Abstract:
Frequently Asked Questions (FAQs) refer to the most common inquiries about specific content. They serve as content comprehension aids by simplifying topics and enhancing understanding through succinct presentation of information. In this paper, we address FAQ generation as a well-defined Natural Language Processing task through the development of an end-to-end system leveraging text-to-text transformation models. We present a literature review covering traditional question-answering systems, highlighting their limitations when applied directly to the FAQ generation task. We propose a system capable of building FAQs from textual content tailored to specific domains, enhancing their accuracy and relevance. We utilise self-curated algorithms to obtain an optimal representation of information to be provided as input and also to rank the question-answer pairs to maximise human comprehension. Qualitative human evaluation showcases the generated FAQs as well-constructed and readable while also utilising domain-specific constructs to highlight domain-based nuances and jargon in the original content.
References:
Alyafeai, Zaid, Maged Saeed AlShaibani, and Irfan Ahmad. 2020. “A Survey on Transfer Learning in Natural Language Processing”. arXiv:2007.04239 [cs.CL]. http://arxiv.org/abs/2007.04239.
Antypas, Dimosthenis, Asahi Ushio, Jose Camacho-Collados, Vitor Silva, Leonardo Neves, and Francesco Barbieri. 2022. “Twitter Topic Classification”. In Proceedings of the 29th International Conference on Computational Linguistics, edited by Nicoletta Calzolari, Chu-Ren Huang, Hansaem Kim, James Pustejovsky, Leo Wanner, Key-Sun Choi, Pum-Mo Ryu, et al., 3386–3400. https://aclanthology.org/2022.coling-1.299.
Chan, Ying-Hong, and Yao-Chung Fan. 2019. “A Recurrent BERT-Based Model for Question Generation”. In Proceedings of the 2nd Workshop on Machine Reading for Question Answering, edited by Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, and Danqi Chen, 154–62. https://doi.org/10.18653/v1/D19-5821.
Das, Bidyut, Mukta Majumder, Santanu Phadikar, and Arif Ahmed Sekh. 2021. “Automatic Question Generation and Answer Assessment: A Survey.” Research and Practice in Technology Enhanced Learning 16 (1). https://doi.org/10.1186/s41039-021-00151-1
Devlin, Jacob, Chang, Ming-Wei, Lee, Kenton, and Toutanova, Kristina. 2019. “BERT: Pretraining of Deep Bidirectional Transformers for Language Understanding”. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4171–4186. https://aclanthology.org/N19-1423/
Du, Xinya and Cardie, Claire. 2018. “Harvesting Paragraph-level Question-Answer Pairs from Wikipedia”. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, 1907–1917. https://aclanthology.org/P18-1177. https://doi.org/10.18653/v1/P18-1177
Du, Xinya, Shao, Junru and Cardie, Claire. 2017. “Learning to Ask: Neural Question Generation for Reading Comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, 1342–1352. https://aclanthology.org/P17-1123. https://doi.org/10.18653/v1/P17-1123
Gunawan, Dani, Sembiring, C.A., and Budiman, Mohammed. 2018. “The Implementation of Cosine Similarity to Calculate Text Relevance between Two Documents”. Journal of Physics: Conference Series, 978, 012120. https://doi.org/10.1088/1742-6596/978/1/012120
Hu, Wenpeng, Liu, Bing, Ma, Jinwen, Zhao, Dongyan, and Yan, Rui. 2018. “Aspect-based Question Generation. 6th International Conference on Learning Representations”. In Workshop Track Proceedings of ICLR 2018, April 30 - May 3. https://openreview.net/forum?id=rkRR1ynIf
Joshi, Mandar, Choi, Eunsol, Weld, Daniel, and Zettlemoyer, Luke. 2017. “TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension”. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, 1601–1611. https://aclanthology.org/P17-1147/. https://doi.org/10.18653/v1/P17-1147
Kay, Anthony. 2007. “Tesseract: an open-source optical character recognition engine”. Linux Journal. 159: 2. https://dl.acm.org/doi/10.5555/1288165.1288167
Kočiský, Tomáš, Schwarz, Jonathan, Blunsom, Phil, Dyer, Chris, M. Hermann, Karl, Melis, Gábor, and Grefenstette, Edward. 2018. “The NarrativeQA Reading Comprehension Challenge”. Transactions of the Association for Computational Linguistics, 6:317–328. https://aclanthology.org/Q18-1023/. https://doi.org/10.1162/tacl_a_00023
Kumar, Vishwajeet, Chaki, Raktim, Talluri, Sai T., Ramakrishnan, Ganesh, Li, Yuan-Fang, and Haffari, Gholamreza. 2019. “Question Generation from Paragraphs: A Tale of Two Hierarchical Models”. arXiv: 1911.03407 [cs.CL]. https://arxiv.org/abs/1911.03407
Kunichika, Hidenobu, Katayama, Tomoki, Hirashima, Tsukasa, and Takeuchi, Akira. 2004.n“Automated question generation methods for intelligent English learning systems and its evaluation”. In Proceedings of ICCE, 2004.
Liu, Bang, Wei, Haojie, Niu, Di, Chen, Haolan, and He, Yancheng. 2020. “Asking Questions the Human Way: Scalable Question-Answer Generation from Text Corpus”. In Proceedings of The Web Conference 2020, 2032–2043. https://doi.org/10.1145/3366423.3380270
Lu, Chao.-Yi, and Lu, Sin-En. 2021. “A Survey of Approaches to Automatic Question Generation: From 2019 to Early 2021”. In Proceedings of the 33rd Conference on Computational Linguistics and Speech Processing (ROCLING 2021), 151–162.mhttps://aclanthology.org/2021.rocling-1.21
Raazaghi, Fatemah. 2015. “Auto-FAQ-Gen: Automatic Frequently Asked QuestionsmGeneration”. In Advances in Artificial Intelligence, edited by Denilson Barbosa and Evangelos Milios, 334–337, Springer International Publishing. https://doi.org/10.1007/978-3-319-18356-5_30
Raffel, Colin, Shazeer, Noam, Roberts, Adam, Lee, Katherine, Narang, Sharan, Matena, Michael, Zhou, Yanqi, Li, Wei, and Liu, Peter J. 2023. “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer”. Journal of Machine Learning Research 21: 1-67. https://jmlr.org/papers/volume21/20-074/20-074.pdf
Rajpurkar, Pranav, Zhang, Jian, Lopyrev, Konstantin, and Liang, Percy. 2016. “SQuAD: 100,000+ Questions for Machine Comprehension of Text”. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, 2383–2392. https://aclanthology.org/D16-1264/. https://doi.org/10.18653/v1/D16-1264
Roemmele, Melissa, Sidhpura, Deep, DeNeefe, Steve, and Tsou, Ling. 2021. “AnswerQuest: A System for Generating Question-Answer Items from Multi-Paragraph Documents”. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, 40–52. https://aclanthology.org/2021.eacl-demos.6/. https://doi.org/10.18653/v1/2021.eacl-demos.6
Ruder, Sebastian, E. Peters, Matthew, Swayamdipta, Swabha and Wolf, Thomas. 2019. “Transfer Learning in Natural Language Processing”. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Tutorials, 15–18. https://aclanthology.org/N19-5004/. https://doi.org/10.18653/v1/N19-5004
Shanthi, M., and Irudhayaraj, Anthony A. 2009. “Multithreading-An Efficient Technique for Enhancing Application Performance”. International Journal of Recent Trends in Engineering, 2 (4).
Shen, Sheng, Yaliang Li, Nan Du, Xian Wu, Yusheng Xie, Shen Ge, Tao Yang, Kai Wang, Xingzheng Liang, and Wei Fan. 2020. “On the Generation of Medical Question-Answer Pairs”. In Proceedings of the AAAI Conference on Artificial Intelligence 34 (05):8822-29. https://doi.org/10.1609/aaai.v34i05.6410.
Trischler, Adam, Wang, Tong, Yuan, Xingdi, Harris, Justin, Sordoni, Alessandro, Bachman, Philip, and Suleman, Kaheer. 2017. “NewsQA: A Machine Comprehension Dataset”. In Proceedings of the 2nd Workshop on Representation Learning for NLP, 191–200. https://aclanthology.org/W17-2623/.https://doi.org/10.18653/v1/W17-2623
Zhang, Shiyue and Bansal, Mohit. 2019. “Addressing Semantic Drift in Question Generation for Semi-Supervised Question Answering”. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2495–2509. https://aclanthology.org/D19-1253/. https://doi.org/10.18653/v1/D19-1253
Zhang, tong, Liu, Yong, Li, Boyang, Zeng, Zhiwei, Wang, Pengwei, You, Yuan, Miao Chunyan, and Cui, Lizhen. 2022. “History-Aware Hierarchical Transformer for Multi-session Opendomain Dialogue System”. In Findings of the Association for Computational Linguistics: EMNLP 2022, 3395–3407. https://aclanthology.org/2022.findings-emnlp.247/
Zhao, Yao, Ni, Xiaochuan, Ding, Yuanyuan, and Ke, Qifa. 2018. “Paragraph-level Neural Question Generation with Maxout Pointer and Gated Self-attention Networks”. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 3901–3910. https://aclanthology.org/D18-1424/.https://doi.org/10.18653/v1/2022.findings-emnlp.247. https://doi.org/10.18653/v1/D18-1424
Zhou, Qingyu, Yang, Nan, Wei, Furu, Tan, Chuanqi, Bao, Hangbo, & Zhou, Ming. 2017. “Neural Question Generation from Text: A Preliminary Study”. arXiv:1704.01792 [cs.CL]. https://arxiv.org/abs/1704.01792



