ERIC - Search Results

Publication Date

In 2025	1
Since 2024	5
Since 2021 (last 5 years)	23

Descriptor

Testing Problems	23
Scores	17
Test Items	9
Foreign Countries	8
Item Response Theory	7
Language Tests	5
Second Language Learning	5
Test Validity	5
Equated Scores	4
Psychometrics	4
Difficulty Level	3
English (Second Language)	3
High Stakes Tests	3
Standardized Tests	3
Student Evaluation	3
Teacher Attitudes	3
Test Reliability	3
Accuracy	2
College Faculty	2
Definitions	2
Effect Size	2
Evaluation Methods	2
Performance	2
Prevention	2
Sample Size	2
More ▼

Publication Type

Journal Articles	20
Reports - Research	16
Reports - Evaluative	4
Dissertations/Theses -…	2
Information Analyses	1
Numerical/Quantitative Data	1
Reports - Descriptive	1

Education Level

Higher Education	4
Postsecondary Education	4
Elementary Education	2
Secondary Education	2
Early Childhood Education	1
Elementary Secondary Education	1
Grade 4	1
Intermediate Grades	1
Preschool Education	1

Audience

Location

United Kingdom	2
China	1
Europe	1
Germany	1
Iran	1
Thailand	1
United States	1

Laws, Policies, & Programs

Every Student Succeeds Act…

Assessments and Surveys

International English…	2
Progress in International…	1

What Works Clearinghouse Rating

Showing 1 to 15 of 23 results Save | Export

Explaining Performance Decline over the Course of Taking Comprehensive Proficiency Tests: The Roles of Effort and Omission Propensity

Peer reviewed

Direct link

Karoline A. Sachse; Sebastian Weirich; Nicole Mahler; Camilla Rjosk – International Journal of Testing, 2024

In order to ensure content validity by covering a broad range of content domains, the testing times of some educational large-scale assessments last up to a total of two hours or more. Performance decline over the course of taking the test has been extensively documented in the literature. It can occur due to increases in the numbers of: (a)…

Descriptors: Test Wiseness, Test Score Decline, Testing Problems, Foreign Countries

Population Invariance in Composite-Score Equating with the Random Groups Design

Direct link

Chang, Kuo-Feng – ProQuest LLC, 2022

This dissertation was designed to foster a deeper understanding of population invariance in the context of composite-score equating and provide practitioners with guidelines for addressing score equity concerns at the composite score level. The purpose of this dissertation was threefold. The first was to compare different composite equating…

Descriptors: Test Items, Equated Scores, Methods, Design

The Development of a Standardized Effect Size for the SIBTEST Procedure

Peer reviewed

Direct link

James D. Weese; Ronna C. Turner; Allison Ames; Xinya Liang; Brandon Crawford – Journal of Experimental Education, 2024

In this study a standardized effect size was created for use with the SIBTEST procedure. Using this standardized effect size, a single set of heuristics was developed that are appropriate for data fitting different item response models (e.g., 2-parameter logistic, 3-parameter logistic). The standardized effect size rescales the raw beta-uni value…

Descriptors: Test Bias, Test Items, Item Response Theory, Effect Size

Estimating Learning When Test Scores Are Missing: The Problem and Two Solutions. EdWorkingPaper No. 23-864

Download full text

Paul T. von Hippel – Annenberg Institute for School Reform at Brown University, 2023

Longitudinal studies can produce biased estimates of learning if children miss tests. In an application to summer learning, we illustrate how missing test scores can create an illusion of large summer learning gaps when true gaps are close to zero. We demonstrate two methods that reduce bias by exploiting the correlations between missing and…

Descriptors: Testing Problems, Scores, Educational Research, Longitudinal Studies

Perceptions of Test Score Pollution Stemming from COVID-19 and State Testing: An Exploratory Case Study

Direct link

Kalemdaroglu-Wheeler, Elif – ProQuest LLC, 2023

The purpose of this qualitative exploratory case study was to explore teachers' and administrators' perceptions of test score pollution deriving from COVID-19-related issues that may affect students' test scores on state-mandated standardized tests for grades six through 12 in a state along the Atlantic Coast of the United States. Four research…

Descriptors: Testing Problems, Scores, COVID-19, Pandemics

Reporting Pass-Fail Decisions to Examinees with Incomplete Data: A Commentary on Feinberg (2021)

Peer reviewed

Direct link

Sinharay, Sandip – Educational Measurement: Issues and Practice, 2022

Administrative problems such as computer malfunction and power outage occasionally lead to missing item scores, and hence to incomplete data, on credentialing tests such as the United States Medical Licensing examination. Feinberg compared four approaches for reporting pass-fail decisions to the examinees with incomplete data on credentialing…

Descriptors: Testing Problems, High Stakes Tests, Credentials, Test Items

Adjusting for Ability Differences of Equating Samples When Randomization Is Suboptimal

Peer reviewed

Direct link

Kim, Sooyeon; Walker, Michael E. – Educational Measurement: Issues and Practice, 2022

Test equating requires collecting data to link the scores from different forms of a test. Problems arise when equating samples are not equivalent and the test forms to be linked share no common items by which to measure or adjust for the group nonequivalence. Using data from five operational test forms, we created five pairs of research forms for…

Descriptors: Ability, Tests, Equated Scores, Testing Problems

To What Degree Does Rapid Guessing Distort Aggregated Test Scores? A Meta-Analytic Investigation

Peer reviewed

Direct link

Rios, Joseph A.; Deng, Jiayi; Ihlenfeldt, Samuel D. – Educational Assessment, 2022

The present meta-analysis sought to quantify the average degree of aggregated test score distortion due to rapid guessing (RG). Included studies group-administered a low-stakes cognitive assessment, identified RG via response times, and reported the rate of examinees engaging in RG, the percentage of RG responses observed, and/or the degree of…

Descriptors: Guessing (Tests), Testing Problems, Scores, Item Response Theory

IRTrees for Skipping Items in PIRLS

Peer reviewed

Direct link

Andrés Christiansen; Rianne Janssen – Educational Assessment, Evaluation and Accountability, 2024

In international large-scale assessments, students may not be compelled to answer every test item: a student can decide to skip a seemingly difficult item or may drop out before the end of the test is reached. The way these missing responses are treated will affect the estimation of the item difficulty and student ability, and ultimately affect…

Descriptors: Test Items, Item Response Theory, Grade 4, International Assessment

Which Assessment Is Harder? Some Limits of Statistical Linking

Download full text

Benton, Tom; Williamson, Joanna – Research Matters, 2022

Equating methods are designed to adjust between alternate versions of assessments targeting the same content at the same level, with the aim that scores from the different versions can be used interchangeably. The statistical processes used in equating have, however, been extended to statistically "link" assessments that differ, such as…

Descriptors: Statistical Analysis, Equated Scores, Definitions, Alternative Assessment

Measurement Invariance of Scores on the Teacher Stress Scale: International Sample of PreK-12 Teachers

Peer reviewed

Direct link

Jiayi Wang; Michael T. Kalkbrenner; Riley Schaner – Psychology in the Schools, 2025

Teaching is a stressful profession with a high turnover rate. Schools and related institutions need to take more action to support teachers and keep teacher stress at a manageable level. The continued research and practical effort require measures to examine teachers' stress in a briefer and accurate manner. The Teacher Stress Scale is a recently…

Descriptors: Elementary School Teachers, Secondary School Teachers, Preschool Teachers, Stress Variables

Item Pool Quality Control in Educational Testing: Change Point Model, Compound Risk, and Sequential Detection

Peer reviewed

Direct link

Chen, Yunxiao; Lee, Yi-Hsuan; Li, Xiaoou – Journal of Educational and Behavioral Statistics, 2022

In standardized educational testing, test items are reused in multiple test administrations. To ensure the validity of test scores, the psychometric properties of items should remain unchanged over time. In this article, we consider the sequential monitoring of test items, in particular, the detection of abrupt changes to their psychometric…

Descriptors: Standardized Tests, Test Items, Test Validity, Scores

Better Remedies for Bad Exams: Correcting for Difficult Questions in a Fair and Systematic Way

Peer reviewed
PDF on ERIC

Download full text

Camenares, Devin – International Journal for the Scholarship of Teaching and Learning, 2022

Balancing assessment of learning outcomes with the expectations of students is a perennial challenge in education. Difficult exams, in which many students perform poorly, exacerbate this problem and can inspire a wide variety of interventions, such as a grading curve. However, addressing poor performance can sometimes distort or inflate grades and…

Descriptors: College Students, Student Evaluation, Tests, Test Items

Assessing Mode Effects of At-Home Testing without a Randomized Trial. Research Report. ETS RR-21-10

Peer reviewed
PDF on ERIC

Download full text

Kim, Sooyeon; Walker, Michael – ETS Research Report Series, 2021

In this investigation, we used real data to assess potential differential effects associated with taking a test in a test center (TC) versus testing at home using remote proctoring (RP). We used a pseudo-equivalent groups (PEG) approach to examine group equivalence at the item level and the total score level. If our assumption holds that the PEG…

Descriptors: Testing, Distance Education, Comparative Analysis, Test Items

Signal-to-Noise Ratio in Estimating and Testing the Mediation Effect: Structural Equation Modeling versus Path Analysis with Weighted Composites

Peer reviewed

Direct link

Ke-Hai Yuan; Zhiyong Zhang; Lijuan Wang – Grantee Submission, 2024

Mediation analysis plays an important role in understanding causal processes in social and behavioral sciences. While path analysis with composite scores was criticized to yield biased parameter estimates when variables contain measurement errors, recent literature has pointed out that the population values of parameters of latent-variable models…

Descriptors: Structural Equation Models, Path Analysis, Weighted Scores, Comparative Testing

Previous Page | Next Page »

Pages: 1 | 2

Educational Measurement:…	2
ProQuest LLC	2
Annenberg Institute for…	1
ETS Research Report Series	1
Educational Assessment	1
Educational Assessment,…	1
Grantee Submission	1
International Journal for the…	1
International Journal of…	1
International Journal of…	1
International Journal on…	1
Journal of Computer Assisted…	1
Journal of Educational and…	1
Journal of Experimental…	1
LEARN Journal: Language…	1
Language Assessment Quarterly	1
Language Testing in Asia	1
Participatory Educational…	1
Psychology in the Schools	1
Research Matters	1
TESOL Quarterly: A Journal…	1
More ▼

Kim, Sooyeon	2
Allison Ames	1
Andrés Christiansen	1
Attali, Yigal	1
Baig, Basim	1
Benton, Tom	1
Brandon Crawford	1
Camenares, Devin	1
Camilla Rjosk	1
Carlsen, Cecilie Hamnes	1
Chanchula, Nawiya	1
Chang, Kuo-Feng	1
Chen, Yunxiao	1
Deng, Jiayi	1
Gu, Kali	1
Horie, André Kenji	1
Hu, Ruolin	1
Ihlenfeldt, Samuel D.	1
James D. Weese	1
Jaturapitakkul, Natjiree	1
Jiayi Wang	1
Kalemdaroglu-Wheeler, Elif	1
Karoline A. Sachse	1
Ke-Hai Yuan	1
Kiliç, Abdullah Faruk	1
More ▼