XÂY DỰNG VÀ THỰC NGHIỆM PHẦN MỀM HỖ TRỢ ĐÁNH GIÁ CHẤT LƯỢNG ĐỀ THI THEO HƯỚNG ĐỊNH LƯỢNG TẠI TRƯỜNG ĐẠI HỌC CÔNG NGHIỆP QUẢNG NINH

DEVELOPMENT AND EXPERIMENTAL IMPLEMENTATION OF SOFTWARE SUPPORTING QUANTITATIVE EXAM QUALITY ASSESSMENT AT QUANG NINH UNIVERSITY OF INDUSTRY

Nguyễn Thu Hiền
Nguyễn Thị Phương
Trương Thị Khánh Ly
Phạm Ngọc Hải

Tóm tắt (Abstract)

Trong bối cảnh các cơ sở giáo dục đại học đang tăng cường đổi mới công tác kiểm tra, đánh giá theo chuẩn đầu ra, việc xây dựng công cụ hỗ trợ đánh giá chất lượng đề thi dựa trên các tiêu chí khoa học và khách quan trở nên cần thiết. Bài báo này trình bày quá trình xây dựng và thực nghiệm phần mềm hỗ trợ đánh giá chất lượng đề thi học phần tại Trường Đại học Công nghiệp Quảng Ninh dựa trên lý thuyết kiểm tra cổ điển (Classical Test Theory – CTT). Phần mềm sử dụng các chỉ số độ khó, độ phân biệt và độ tin cậy Cronbach’s Alpha; trong đó, bách phân vị được sử dụng để xác định nhóm người học có kết quả cao và nhóm người học có kết quả thấp khi tính độ phân biệt. Phần mềm được xây dựng bằng ngôn ngữ Python, sử dụng thư viện Pandas và NumPy để xử lý dữ liệu, Streamlit để xây dựng giao diện người dùng. Phần mềm hỗ trợ kiểm tra và chuẩn hoá dữ liệu đầu vào, tính toán các chỉ số, phân loại câu hỏi, cảnh báo các trường hợp cần rà soát và xuất báo cáo kết quả. Qua đó, phần mềm góp phần nâng cao hiệu quả công tác khảo thí và cung cấp căn cứ phục vụ đánh giá mức độ đạt chuẩn đầu ra học phần tại Trường Đại học Công nghiệp Quảng Ninh. Phần mềm được thử nghiệm trên 41 đề thi thuộc 29 học phần của các chương trình đào tạo khác nhau trong học kỳ I năm học 2025-2026, gồm 04 đề trắc nghiệm, 19 đề tự luận và 18 đề hỗn hợp. Kết quả thực nghiệm cho thấy phần mềm có khả năng xử lý cả 03 hình thức đề thi, tự động tính toán các chỉ số và cung cấp thông tin phục vụ việc nhận diện câu hỏi quá dễ, quá khó, có độ phân biệt thấp hoặc có dấu hiệu bất thường. Tuy nhiên, với lớp có quy mô nhỏ hơn 10 sinh viên, kết quả độ phân biệt chỉ nên được xem là chỉ báo tham khảo. Việc kết luận câu hỏi giữ lại, chỉnh sửa hay loại bỏ cần kết hợp các chỉ số về độ khó, độ phân biệt, độ tin cậy và nhận định của giảng viên và bộ môn.
In the context of higher education institutions increasingly reforming assessment practices in alignment with learning outcomes, developing a tool to support the evaluation of examination quality based on scientific and objective criteria has become necessary. This paper presents the development and experimental implementation of software designed to support the evaluation of course examination quality at Quang Ninh University of Industry based on Classical Test Theory (CTT). The software employs difficulty, discrimination, and Cronbach’s Alpha reliability indices; percentile ranking is used to identify high-performing and low-performing groups when calculating the discrimination index. The software was developed in Python, using Pandas and NumPy for data processing and Streamlit for the user interface. The software supports the validation and standardization of input data, calculation of indices, classification of test items, identification of cases requiring review, and generation of analytical reports. Through these functions, the software contributes to improving the effectiveness of assessment activities and provides evidence to support the evaluation of course learning outcomes at Quang Ninh University of Industry. The software was tested on 41 examinations from 29 courses across different academic programs in the first semester of the 2025–2026 academic year, including 4 multiple-choice examinations, 19 essay examinations, and 18 mixed-format examinations. The experimental results show that the software can process all three examination formats, automatically calculate the relevant indices, and provide information for identifying items that are overly easy, overly difficult, have low discrimination, or show abnormal patterns. However, for classes with fewer than 10 students, the discrimination index should be regarded only as a reference indicator. Decisions on whether to retain, revise, or remove an item should be based on a combined consideration of difficulty, discrimination, reliability, and the professional judgment of lecturers and academic departments.
Bạn đã không sử dụng Site, Bấm vào đây để duy trì trạng thái đăng nhập. Thời gian chờ: 60 giây
Gửi phản hồi