Publication Date
In 2025 | 10 |
Since 2024 | 22 |
Since 2021 (last 5 years) | 86 |
Since 2016 (last 10 years) | 168 |
Since 2006 (last 20 years) | 249 |
Descriptor
Difficulty Level | 249 |
Test Validity | 249 |
Test Items | 157 |
Test Reliability | 137 |
Foreign Countries | 122 |
Test Construction | 96 |
Item Response Theory | 54 |
Psychometrics | 48 |
Multiple Choice Tests | 43 |
Item Analysis | 42 |
Scores | 39 |
More ▼ |
Source
Author
Liu, Kimy | 3 |
Paek, Insu | 3 |
Schoen, Robert C. | 3 |
Tindal, Gerald | 3 |
Yang, Xiaotong | 3 |
Alexander, Patricia A. | 2 |
Baghaei, Purya | 2 |
Beege, Maik | 2 |
Bejar, Isaac I. | 2 |
Chen, Jing | 2 |
Crisp, Victoria | 2 |
More ▼ |
Publication Type
Education Level
Audience
Administrators | 1 |
Community | 1 |
Parents | 1 |
Teachers | 1 |
Location
Turkey | 15 |
Indonesia | 13 |
Germany | 7 |
Iran | 6 |
Japan | 6 |
Nigeria | 6 |
United Kingdom | 4 |
California | 3 |
Canada | 3 |
Chile | 3 |
China | 3 |
More ▼ |
Laws, Policies, & Programs
Pell Grant Program | 1 |
Assessments and Surveys
What Works Clearinghouse Rating
Camilo Vieira; Andrea Vásquez; Federico Meza; Roxana Quintero-Manes; Pedro Godoy – ACM Transactions on Computing Education, 2024
Currently, there is little evidence about how non-English-speaking students learn computer programming. For example, there are few validated assessment instruments to measure the development of programming skills, especially for the Spanish-speaking population. Having valid assessment instruments is essential to identify the difficulties of the…
Descriptors: Programming, Spanish Speaking, Translation, Test Validity
Sherwin E. Balbuena – Online Submission, 2024
This study introduces a new chi-square test statistic for testing the equality of response frequencies among distracters in multiple-choice tests. The formula uses the information from the number of correct answers and wrong answers, which becomes the basis of calculating the expected values of response frequencies per distracter. The method was…
Descriptors: Multiple Choice Tests, Statistics, Test Validity, Testing
Tia M. Fechter; Heeyeon Yoon – Language Testing, 2024
This study evaluated the efficacy of two proposed methods in an operational standard-setting study conducted for a high-stakes language proficiency test of the U.S. government. The goal was to seek low-cost modifications to the existing Yes/No Angoff method to increase the validity and reliability of the recommended cut scores using a convergent…
Descriptors: Standard Setting, Language Proficiency, Language Tests, Evaluation Methods
Krieglstein, Felix; Beege, Maik; Rey, Günter Daniel; Sanchez-Stockhammer, Christina; Schneider, Sascha – Educational Psychology Review, 2023
According to cognitive load theory, learning can only be successful when instructional materials and procedures are designed in accordance with human cognitive architecture. In this context, one of the biggest challenges is the accurate measurement of the different cognitive load types as these are associated with various activities during…
Descriptors: Test Construction, Test Validity, Questionnaires, Cognitive Processes
Hojung Kim; Changkyung Song; Jiyoung Kim; Hyeyun Jeong; Jisoo Park – Language Testing in Asia, 2024
This study presents a modified version of the Korean Elicited Imitation (EI) test, designed to resemble natural spoken language, and validates its reliability as a measure of proficiency. The study assesses the correlation between average test scores and Test of Proficiency in Korean (TOPIK) levels, examining score distributions among beginner,…
Descriptors: Korean, Test Validity, Test Reliability, Imitation
Jerin Kim; Kent McIntosh – Journal of Positive Behavior Interventions, 2025
We aimed to identify empirically valid cut scores on the positive behavioral interventions and supports (PBIS) Tiered Fidelity Inventory (TFI) through an expert panel process known as bookmarking. The TFI is a measurement tool to evaluate the fidelity of implementation of PBIS. In the bookmark method, experts reviewed all TFI items and item scores…
Descriptors: Positive Behavior Supports, Cutting Scores, Fidelity, Program Evaluation
Krieglstein, Felix; Beege, Maik; Rey, Günter Daniel; Ginns, Paul; Krell, Moritz; Schneider, Sascha – Educational Psychology Review, 2022
For more than three decades, cognitive load theory has been addressing learning from a cognitive perspective. Based on this instructional theory, design recommendations and principles have been derived to manage the load on working memory while learning. The increasing attention paid to cognitive load theory in educational science quickly…
Descriptors: Cognitive Processes, Difficulty Level, Learning Theories, Test Reliability
Aditya Shah; Ajay Devmane; Mehul Ranka; Prathamesh Churi – Education and Information Technologies, 2024
Online learning has grown due to the advancement of technology and flexibility. Online examinations measure students' knowledge and skills. Traditional question papers include inconsistent difficulty levels, arbitrary question allocations, and poor grading. The suggested model calibrates question paper difficulty based on student performance to…
Descriptors: Computer Assisted Testing, Difficulty Level, Grading, Test Construction
Ober, Teresa M.; Lu, Yikai; Blacklock, Chessley B.; Liu, Cheng; Cheng, Ying – Journal of Psychoeducational Assessment, 2023
We develop and validate a self-report measure of intrinsic and extrinsic cognitive load suitable for measuring the constructs in a variety of learning contexts. Data were collected from three independent samples of college students in the U.S. (N[subscript total]= 513; M[subscript age]= 21.13 years). Kane's (2013) framework was used to validate…
Descriptors: Test Construction, Test Validity, Cognitive Processes, Difficulty Level
Menold, Natalja; Raykov, Tenko – Educational and Psychological Measurement, 2022
The possible dependency of criterion validity on item formulation in a multicomponent measuring instrument is examined. The discussion is concerned with evaluation of the differences in criterion validity between two or more groups (populations/subpopulations) that have been administered instruments with items having differently formulated item…
Descriptors: Test Items, Measures (Individuals), Test Validity, Difficulty Level
Miller, Dan J.; Noble, Prisca; Medlen, Sue; Jones, Karina; Munns, Suzanne L. – Journal of Experimental Education, 2023
The cognitive load imposed by instruction is an important consideration for instructional designers. Theoretical models have traditionally divided total cognitive load into intrinsic, extrinsic, and germane load. The 10-item Cognitive Load Inventory (CLI-10) is designed to measure these three types of cognitive load. It is typically administered…
Descriptors: Psychometrics, Cognitive Processes, Difficulty Level, Factor Analysis
Apichat Khamboonruang – Language Testing in Asia, 2025
Chulalongkorn University Language Institute (CULI) test was developed as a local standardised test of English for professional and international communication. To ensure that the CULI test fulfils its intended purposes, this study employed Kane's argument-based validation and Rasch measurement approaches to construct the validity argument for the…
Descriptors: Universities, Second Language Learning, Second Language Instruction, Language Tests
Suwita Suwita; Sulistyo Saputro; Sajidan Sajidan; Sutarno Sutarno – Journal of Baltic Science Education, 2024
The current study uses the Rasch Model to measure lower-secondary school students' critical thinking skills on photosynthesis topics. Critical thinking skills are considered essential in science education, but few valid and practical measurement instruments remain. The current study fills the gap by adapting the instrument from the Watson-Glaser…
Descriptors: Secondary School Students, Critical Thinking, Thinking Skills, Botany
Douglas-Morris, Jan; Ritchie, Helen; Willis, Catherine; Reed, Darren – Anatomical Sciences Education, 2021
Multiple-choice (MC) anatomy "spot-tests" (identification-based assessments on tagged cadaveric specimens) offer a practical alternative to traditional free-response (FR) spot-tests. Conversion of the two spot-tests in an upper limb musculoskeletal anatomy unit of study from FR to a novel MC format, where one of five tagged structures on…
Descriptors: Multiple Choice Tests, Anatomy, Test Reliability, Difficulty Level
Ruying Li; Gaofeng Li – International Journal of Science and Mathematics Education, 2025
Systems thinking (ST) is an essential competence for future life and biology learning. Appropriate assessment is critical for collecting sufficient information to develop ST in biology education. This research offers an ST framework based on a comprehensive understanding of biological systems, encompassing four skills across three complexity…
Descriptors: Test Construction, Test Validity, Science Tests, Cognitive Tests