https://doi.org/10.67147/literariness.v1si1.010
Comparative Evaluation of AI Chatbot Output Quality Using NLP Metrics and Human Expert Assessment
S. DHANUSH
Research Scholar
Department of English
Vels Institute of Science, Technology and Advanced Studies, Pallavaram, Chennai, India.
DR. P. SANTHOSH
Assistant Professor
Department of English
Vels Institute of Science, Technology and Advanced Studies, Pallavaram, Chennai
Abstract: As long as human society continues to evolve, technological advancement will remain inevitable. In the contemporary era, the focus has shifted from merely possessing technology to understanding and controlling the systems that drive it. Generative AI chatbots such as ChatGPT, Google Gemini, Perplexity, Claude, and Grok have become integral tools in various aspects of daily life, including education, research, and professional communication.
Existing studies have primarily examined individual chatbot models, focusing on the quality, accuracy, and contextual relevance of their responses. This study undertakes a comparative evaluation of major AI chatbots by providing them with twenty standardized prompts and collecting their responses for analysis. The research adopts a mixed-methods approach, combining automated Natural Language Processing (NLP) metrics with structured human expert assessment.
The quantitative analysis evaluates response quality using NLP metrics such as BLEU, ROUGE, BERTScore, and grammar analysis to measure accuracy, semantic similarity, and linguistic correctness. The qualitative component involves expert evaluation of the responses based on criteria including factual accuracy, grammatical quality, content completeness, coherence, and contextual relevance. The findings are compared to identify variations in the output quality of different AI chatbots. This study aims to contribute to a deeper understanding of the strengths and limitations of contemporary AI chatbots and their effectiveness in academic and knowledge-based contexts.
Keywords: Artificial Intelligence, Chatbot Evaluation, ChatGPT, Google Gemini, Perplexity, Claude, Grok, Natural Language Processing, Human Expert Assessment, Generative AI
Read full paper PDF