Програмний модуль автоматизованого збору та оцінки відповідей LLM-моделі
Loading...
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Тернопіль, ЗУНУ
Abstract
Метою кваліфікаційної роботи є розробка програмного модуля автоматизованого збору, збереження та оцінювання відповідей великих мовних моделей на основі сучасних методів семантичного аналізу тексту та підходу LLM-asa-Judge.
Об’єктом дослідження є процес автоматизованої взаємодії з великими мовними моделями та аналіз якості сформованих ними відповідей.
Предметом дослідження є методи, алгоритми та програмні засоби збору, збереження, порівняння й оцінювання відповідей LLM-моделей.
У результаті виконання роботи розроблено програмний модуль, який забезпечує автоматизоване надсилання тестових запитів до великих мовних моделей через API, збереження отриманих відповідей у базі даних та їх подальше оцінювання за допомогою метрик семантичної подібності. Для підвищення об’єктивності оцінювання реалізовано механізм LLM-as-a-Judge, що дозволяє використовувати окрему мовну модель у ролі експертного оцінювача якості відповідей.
Розроблена система може бути використана для тестування та порівняння сучасних мовних моделей, проведення бенчмаркінгу AI-сервісів, аналізу ефективності промптів та підтримки досліджень у галузі штучного інтелекту й обробки природної мови.
The purpose of the qualification work is to develop a software module for automated collection, storage, and evaluation of responses generated by Large Language Models (LLMs) using modern semantic text analysis methods and the LLM-as-a-Judge approach.
The object of research is the process of automated interaction with large language models and the analysis of the quality of generated responses.
The subject of research is methods, algorithms, and software tools for collecting, storing, comparing, and evaluating responses produced by LLM-based systems.
As a result of the research, a software module was developed that provides automated submission of test queries to large language models through API interfaces, storage of generated responses in a database, and their further evaluation using semantic similarity metrics. To improve the objectivity of the evaluation process, the LLM-as-a-Judge approach was implemented, allowing an additional language model to perform the role of an expert evaluator.
The developed system can be used for benchmarking modern language models, evaluating AI services, testing prompt engineering strategies, and supporting research in the fields of Artificial Intelligence and Natural Language Processing. The modular architecture of the system enables easy integration of new language models and evaluation methods without significant modifications to the existing software structure.
Description
Citation
Копилов, М. О. Програмний модуль автоматизованого збору та оцінки відповідей LLM-моделі = Software Module for Automated Collection and Evaluation of LLM Model Responses : кваліфікаційна робота : спец. 122 – комп’ютерні науки ; освітньо-професійна програма – комп’ютерні науки / Максим Олександрович Копилов ; наук. керівник к.т.н., доц. П. Є. Биковий. Тернопіль : ЗУНУ, 2026. 70 с.