Do interviews change participants’ behaviors?: Using log data to investigate student reactivity during data-driven classroom interview
Hyeongjo Kim1, Luc Paquette1, Yiqiu Zhou1, Qianhui Liu1, Jaclyn Ocumpaugh2, Andres Felipe Zambrano3, Amanda Barany3
1University of Illinois at Urbana-Champaign
2University of Houston
3University of Pennsylvania
{hk61, yiqiuz3, ql29}@illinois.edu,
{luc.paquette, jlocumpaugh, afzambrano97, amanda.barany}@gmail.com
Do not delete, move, or resize this block. If the paper is accepted, this block will need to be filled in with reference information.

ABSTRACT

Data-Driven Classroom Interview (DDCI) uses real-time student data to identify important learning events and support the collection of rich in-situ interviews at critical moments of the learning process. Because DDCI occurs during the students' use of the learning system, they may introduce reactivity, whereby participation in the research may alter the experience of the learners. This study investigates short-term behavioral reactivity associated with DDCI and examines a potential mechanism—the attention cue effect—within What-If Hypothetical Implementations in Minecraft (WHIMC), a game-based learning environment. We used log data to compare the learning behaviors of interviewed and non-interviewed students on a set of key WHIMC behaviors, including movement, scientific observation, and measurement tool use. Results show that DDCI is associated with significant behavioral reactivity for behaviors that are more intentional and goal directed (student in-game observations and tool usage) but not for a third behavior (movement), which may function as less direct, peripheral byproduct of task engagement. Moreover, reactivity was not explained by explicit discussions of these behaviors during DDCI, which may indicate that reactivity is arising through other mechanisms. By empirically documenting short-term behavioral reactivity in DDCI and systematically examining its potential mechanisms, this study contributes to the methodological refinement of DDCI and highlights important considerations for the design and interpretation of data-driven research in open-ended learning environments.

Keywords

Data-driven classroom interview (DDCI), In-situ interviewing, Reactivity, Behavior log data analysis, Game-based learning

INTRODUCTION

Data-Driven Classroom Interview (DDCI) is a research method designed to deepen our understanding of learning phenomena by complementing traditional in-situ interviews using detectors based on learner modeling methods [1, 20]. While in-situ interviews have been used to capture learners’ experiences and meanings within authentic learning contexts, they face two persistent methodological challenges [1]. First, researchers must determine how to identify rare yet pedagogically important events that are difficult to observe directly, often described as a “needle-in-a-haystack" ” problem [1]. Second, in-situ interviews raise a sampling issue [22, 25], namely, deciding whom to interview and when, without relying solely on convenience or researcher intuition. By operationalizing theoretically and empirically meaningful learning behaviors through log data to inform interview opportunities, DDCI enables researchers to capture critical yet otherwise elusive learning events and to interview participants based on principled criteria [20].

DDCI belongs to a range of in situ methods--including more traditional interviewing techniques, think-aloud protocols [8], and ecological momentary assessment (EMA; [19])--that can improve validity by allowing the researcher access to the students thoughts about the learning experience as it unfolds. However, just like other in-situ research methods, DDCI alters the learning environment for at least the duration of that research activity, and possibly beyond. These measurement effects have a range of labels in the literature including the observer effect [30] or reactivity [4, 34]. Such effects may stem from increased awareness of being observed (the Hawthorne effect; [18]), motivational responses to perceived evaluation or comparison (the John Henry effect; [27]), or attempts to align with perceived researcher expectations (the observer-expectancy effect; [26]). Whereas earlier methodological traditions often sought to eliminate or control reactivity, recent scholarship increasingly emphasizes explicitly investigating and articulating reactivity as part of the research process [17, 34].

Several considerations are particularly important for investigating reactivity in the context of DDCI. First, reactivity research should examine behaviors that are directly logged by the system and that simultaneously relate to learning processes. Reactivity may manifest across multiple domains, including emotion, behavior, and metacognition. Prior work in educational contexts has explored some of these dimensions. For example, Liu et al. [16] found that observation influenced students’ display of emotional states by increasing engaged concentration while reducing frustration and off-task behavior. Similarly, Bosch et al.’s [3] analysis of metacognitive content in DDCI found that certain forms of verbalized metacognition were associated with learning outcomes, particularly for students who engaged in metacognitive verbalization during interviews. However, these studies did not directly examine behavioral changes in the log data that serve as the triggering events for DDCI and may also represent important learning behaviors. This omission is consequential, as changes in logged behavior may alter detector activation, subsequent interview selection, and, ultimately, interpretations of learning.

In addition, systematic approaches to studying reactivity should be considered. Reactivity can be more rigorously examined by (1) comparing participants’ behaviors before and after the measurement of interest, thereby leveraging temporal changes associated with repeated measurement [10], and (2) employing between-participant comparisons between students who are measured and those who are not, in order to distinguish reactivity from other confounding factors. In the DDCI context, prior logged behaviors may determine which students are interviewed, and reactivity may be intertwined with both interview-related influences and observer effects. Therefore, research on DDCI reactivity should consider both students’ behaviors before interviews and the behaviors of non-interviewed peers to control the observer effects.

Finally, understanding the mechanisms underlying reactivity is crucial for explaining why reactivity occurs and how it can be mitigated in DDCI settings. Several mechanisms may plausibly explain DDCI-related reactivity, including observer effects, verbalization effects, and metacognitive activation. The presence of researchers in the classroom may elicit observer effects, as evidenced by prior findings [16, 30]. Additionally, DDCI interviews require students to engage in high-level verbalization about their thoughts and actions, which is likely to induce stronger reactivity than lower-level verbalization methods such as think-aloud protocols [6, 8]. Such verbalization may function as an attentional cue [4] that stimulates cognitive and metacognitive processes, which can vary depending on both individual characteristics (e.g., interest) and task features (e.g., writing versus reading). Given these multiple plausible mechanisms, researchers should design studies and analyses that aim to identify which mechanisms most plausibly account for observed reactivity effects.

The present study examines whether and how students’ behaviors are influenced by DDCI in WHIMC (What-if Hypothetical Implementations in Mine-craft), a Minecraft-based learning environment [9, 14]. Specifically, we investigate reactivity in behaviors that are directly logged by the system by comparing the behaviors of interviewed students to the behavior of peers who were not interviewed during the same activity. When interview-related reactivity is identified, we further examine whether such effects are associated with direct attentional cues, through which explicit references to specific behaviors (e.g., “observation”) during the interview may promote subsequent enactment of those behaviors. Based on the observed patterns of reactivity, we discuss potential distortions in DDCI-based analyses. By clarifying the nature and implications of behavioral reactivity in DDCI, this study advances the methodological development and supports the careful implementation of data-driven interviews.

RELATED WORK

Data-Driven Classroom Interview

Data-Driven Classroom Interview (DDCI; [1, 20]) is an emerging research method designed to deepen understanding of learning phenomena by augmenting traditional in-situ interviews with detector technologies. By operationalizing focal learning behaviors through log data, detectors can automatically identify moments when targeted behaviors occur and trigger interviews accordingly. This approach allows researchers to capture critical but otherwise elusive learning events and to interview participants based on theoretically and empirically meaningful criteria. The feasibility of such detectors has been substantially enhanced by advances in learning analytics (LA) and educational data mining (EDM), which have demonstrated that a wide range of learning phenomena can be defined, modeled, and monitored using rich streams of behavioral data, such as interaction logs [29].

The DDCI process typically proceeds as follows [20]. Based on the research questions, researchers first develop behavioral models that define interview triggers—specific patterns of activity that research hopes to investigate. Researchers can assign priorities to different triggers to detect rare or theoretically important behaviors while also distributing interview opportunities across participants. When the DDCI system detects a predefined behavior, it alerts researchers, who can then decide whether and when to conduct an interview. These interviews are typically brief (often under five minutes) to minimize disruption to learners’ ongoing activity while capturing highly contextualized and ecologically valid accounts of in-situ experiences. In summary, the core stages of DDCI include trigger development, priority setting, and interviewer engagement following trigger detection.

Because DDCI relies on data-driven technologies, their systematic implementation requires technical infrastructure [20]; triggers must be developed and integrated with the learning environment’s logging system to facilitate real-time detection and notification as well as the recording and storage of interview data. To date, this has been conducted with Quick Red Fox (QRF), an open-source server–client Android application that streamlines in-situ data collection by directing interviewers to learners immediately after the occurrence of research-relevant behaviors defined a priori by the research team [20].

By integrating data-driven detection with in-situ interviewing, DDCI offers methodological benefits for both qualitative and quantitative research. While in-situ interviews have been used to capture learners’ experiences and meanings within authentic learning contexts [13], they face two persistent methodological challenges [1]. First, researchers must identify rare yet pedagogically significant events that are difficult to observe directly—a challenge often characterized as a “needle-in-a-haystack” problem. Second, in-situ interviews raise a sampling issue [22, 25]: deciding whom to interview and when, without relying primarily on convenience or researcher intuition. DDCI addresses these challenges by leveraging detectors derived from log data.

For quantitative research, behavioral data alone provide limited access to the contextual meanings and latent variables underlying observable actions, which raises concerns related to construct validity and inferential plausibility [32]. Qualitative data such as DDCI can complement purely quantitative approaches by addressing challenges in interpreting behavioral data, particularly when inferring learners’ intentions or reasoning from logs alone [1, 11]. DDCI provides learners with opportunities to reflect on and explain their behaviors, generating insights that can inform feature interpretation, model refinement, and theory development.

However, the advantages of in-situ interviews are accompanied by an important methodological concern. Because interviews occur during ongoing learning activities in DDCI, participants may unintentionally alter their behavior in response to being interviewed or to their awareness of research participation. Understanding and accounting for such changes is therefore essential when interpreting data collected through DDCI.

Reactivity

Reactivity refers to any changes that arise from being measured or participating in research itself, with potential effects on participants’ motivation, expectations, affect, and goals [4]. Empirical work has documented reactivity across a range of research methods, including ecological momentary assessment (EMA; [19]), think-aloud protocols [6, 8, 28], and interview-based approaches [3, 23, 24]. In a broad sense, all forms of research participation entail some degree of influence on participants. In any interview-based method, such as DDCI, allocating time for an interview and responding to researchers’ questions are intended and necessary influence as a part of data collection process. Therefore, it is important to delimit the scope of reactivity under consideration. Accordingly, the form of reactivity most relevant to the present study is what Zahle [34] terms ‘regularly unintended reactivity'—that is, routinely occurring effects of research participation that extend beyond the method’s intended influence.

Traditional methodological perspectives, particularly within experimental psychology, have emphasized minimizing or eliminating reactivity [15]. From this standpoint, reactivity is treated primarily as a threat to validity. For instance, introspection-based methods, which require participants to report on their internal cognitive processes through surveys or interviews, have been especially criticized on these grounds. In response, proponents of think-aloud protocols developed theoretical frameworks distinguishing levels of verbalization, arguing that low-level verbalizations—such as articulating information already in working memory—introduce minimal reactivity and can therefore serve as valid data sources [6].

In contrast, other research traditions—particularly qualitative inquiry—have taken a more generative view of participant influence. Rather than framing interview effects as reactivity to be controlled, qualitative researchers often conceptualize them as reflexivity [23]. From this perspective, interviews are understood as opportunities for participants to reflect on, reinterpret, and reorganize their experiences in interaction with the researcher. Prior work has examined both during-interview and post-interview reflexivity, highlighting how participation in research can shape participants’ self-understanding over time [23]. This orientation reflects an epistemological stance in which interviews are not treated as neutral pipelines for extracting pre-existing meanings, but as sites of meaning and knowledge construction through interviewer–interviewee interaction [13]. Within this framework, participant influence is not a methodological flaw to be eliminated but a natural and often productive feature of the research process.

More recent methodological work seeks to bridge these perspectives by advocating for a reactivity-transparent approach, in which reactivity is explicitly examined, documented, and analytically incorporated rather than treated solely as bias [17, 34]. Zahle [34] argues that researchers should supplement data with explicit assumptions about reactivity, enabling principled reasoning about how collected data may have been shaped by the research process. Depending on the assumptions adopted, researchers may (a) claim reactivity has been adequately controlled, (b) acknowledge partial reactivity effects, or (c) actively investigate the sources and forms of reactivity to infer how participants might have behaved in the absence of measurement.

Strategies have been proposed for studying reactivity itself rather than treating it exclusively as a source of contamination. Gieles et al. [10], for example, distinguish between direct and indirect approaches to assessing measurement reactivity. Direct approaches examine reactivity by explicitly asking participants about perceived changes resulting from research participation, typically through interviews or self-report questionnaires. These methods provide insight into participants’ own interpretations of how measurement influenced their thinking, awareness, or behavior, but they are limited by retrospection, social desirability, and participants’ ability to articulate subtle or implicit changes.

Indirect approaches [10], by contrast, examine reactivity by detecting systematic changes in responses over repeated measurement, rather than relying on participants’ self-reports about influence. Such approaches are especially relevant in EMA contexts, where repeated self-monitoring may promote reflection or learning. For instance, researchers can compare whether survey-based self-perceptions align more strongly with EMA-recorded experiences after, rather than before, an EMA period to assess how EMA may enhance self-understanding (i.e., reactivity in this context). Patterns of increasing alignment or shifts across measurement phases may indicate that participation in measurement itself altered participants’ awareness or behavior. These frameworks provide a foundation for examining reactivity in DDCI, where repeated interviews are conducted in real time.

Reactivity of Data-Driven Classroom Interview

Because DDCI, like other in-situ methodologies, may introduce multiple sources of reactivity, examining its reactivity is essential to understand and articulate how it affects students’ learning experiences. First, DDCI requires researchers to be physically present in the classroom, which may produce observer effects [16, 30]. The visible presence of researchers can heighten students’ awareness of being observed and lead them to modify their behavior relative to their typical classroom practices. For example, Liu et al. [16] compared students’ emotional states in classroom settings with and without observers and found that researcher observation alone could differentially affect detected emotions—specifically increasing engaged concentration while reducing frustration or off-task behaviors. Such effects may be partly attributable to implicit power asymmetries between interviewers and students. In classroom contexts, interviewers are often perceived as authority figures or evaluators, which may inhibit students from sharing authentic accounts of their experiences [1, 31]. To alleviate the issue, DDCI prioritizes eliciting learners’ accounts of their ongoing activities and experiences from their own perspectives [1, 20].

Reactivity in DDCI may also vary depending on the level of verbalization elicited during interviews and the interview strategies used to prompt it. Ericsson and Simon [5, 7] distinguished three levels of verbalization. Level 1 verbalization involves the direct vocalization of inner speech as it occurs, whereas Level 2 verbalization requires verbal encoding and organization of thoughts already present in working memory. Although both levels may slow or extend cognitive processing due to overt verbalization, they are generally assumed not to alter the underlying sequence of thought [6, 8]. In contrast, Level 3 verbalization—often referred to as introspection—requires participants to explain or justify their thinking (e.g., “Why are you solving the problem this way?”). This level introduces additional cognitive demands that can interrupt ongoing cognitive processes and resume them in a qualitatively different state. From this perspective, reactivity in DDCI is not uniform but might depend critically on interview design and the extent to which interview prompts elicit higher levels of reflective or explanatory verbalization.

Empirical evidence supports this distinction. In DDCI contexts where interviewers frequently prompted students to articulate their problem-solving strategies, Bosch et al. [3] found that interviews were associated with increased learning outcomes, particularly for students who engaged in metacognitive verbalization. These findings suggest that reactivity in DDCI may arise not only from the act of interviewing itself but also from the specific content and cognitive demands of interview or its prompts. Consequently, both interview presence and interview design must be considered when examining DDCI reactivity.

The impact of DDCI can be further understood through metacognitive mechanisms activated by interview content. Double and Birney [4] proposed a theoretical framework that conceptualizes reactivity as arising from changes in the cognitive–metacognitive system induced by measurement. Drawing primarily on research on metacognitive measures such as think-aloud protocols, confidence judgments, and judgments of learning, their framework suggests that measurement directs learners’ attention toward particular diagnostic cues—either about their ongoing experience (e.g., perceived difficulty, uncertainty, effort) or about their self-beliefs (e.g., perceived competence or memory ability). Attending these cues influences metacognitive monitoring, which in turn shapes metacognitive control decisions and ultimately affects cognitive performance [4, 12]. These cues function as inputs for diagnosing one’s performance or motivation relative to task goals, thereby influencing decisions about strategy selection, effort allocation, and persistence.

Importantly, as Double & Birney [4] showed, the strength and direction of such reactivity depend on multiple interacting factors, including task characteristics (e.g., complexity, novelty, openness) and person characteristics (e.g., prior knowledge, metacognitive skill, motivation). These factors shape how measurement-induced cues are interpreted and acted upon. Accordingly, within this metacognitive framework, reactivity is not viewed as random noise, but as a systematic and theoretically explainable consequence of how measurement interacts with learners’ monitoring and control processes. This perspective provides a principled basis for examining reactivity in DDCI and for interpreting its effects on learners’ behavior and learning trajectories.

METHODS

WHIMC

WHIMC (What-If Hypothetical Implementation in Minecraft) is a suite of Minecraft-based simulations designed to support astronomy-focused inquiry in order to foster student interest [9, 14]. The WHIMC environment includes thirteen space simulation worlds: two orientation worlds (Rocket Launch and Space Station Hub), six “what-if” Earth scenarios (No Moon, Colder Sun, Tilted Earth, Two Moons, and Earth as Moon), five exoplanet worlds (Trappist, Kepler, Gliese, Cancri, and Brown Dwarf), and a Mars world centered on habitat construction.

Within WHIMC, students explore hypothetical questions by comparing normal Earth to simulated alternative conditions, such as an Earth without the Moon or with a colder Sun. As they navigate these worlds, learners freely explore the environment, identify distinctive features at points of interest (e.g., shorter vegetation or the presence of wind turbines), and use embedded scientific tools (e.g., wind-speed sensors) to make observations and collect data. These interactions encourage students to pose questions or generate hypotheses about observed phenomena (e.g., why plants are shorter in the absence of the Moon or whether stronger winds might be responsible). Figure 1 presents example screen captures of gameplay in WHIMC.

Screenshot from WHIMC, a minecraft-based game world, showing a wind turbine on a hill. On-screen text displays measured airflow speed, and a student asks why the wind is so fast.
Figure 1. Example of gameplay in WHIMC. Students examine a wind turbine in the Earth without Moon scenario and learn that it serves as a primary source of power generation. They can then measure wind speed and record an observation by asking questions or making inferences (e.g., “Why is the wind so fast?”).

To investigate their questions, students seek evidence through multiple forms of in-game scaffolding, including examining environmental features at points of interest, using scientific tools to gather measurements, and interacting with non-playable characters (NPCs) whose quests and floating hints guide inquiry. As students connect observations to data and explanations, their inquiry can progress toward deeper causal understanding of the underlying “what-if” scenario—for example, linking increased wind speeds to faster Earth rotation and ultimately to the Moon’s gravitational influence. At the same time, the open-ended design of WHIMC means that inquiry is not strictly linear: students retain substantial autonomy and may pursue multiple exploratory paths as they navigate the simulated worlds.

Data-Driven Classroom Interviewing

Participants

This study examined data collected from five-day science camps for middle school students that employed WHIMC. On the first day of the camp, participants completed surveys assessing individual interest, prior knowledge, and prior game experience. During the first three days, students explored a series of “what-if” simulation worlds and investigated the distinctive characteristics of each environment. At the conclusion of Day 3, post-surveys measuring individual interest, situational interest, and knowledge were administered. During the final two days of the camp, students engaged in a Mars-based habitat construction activity. Throughout the whole gameplay, Quick Red Fox (QRF), an Android application for conducting DDCI [20], was used to detect predefined behavioral patterns in students’ in-game actions (e.g., repeated use of the same scientific tool) and to identify candidates for in-situ interviews.

The dataset analyzed in this study combined data from four camps conducted in different regions of the United States, including the Midwest and Mountain West, during 2024 and 2025. The sample consisted of 98 participants, of whom 44 identified as male, 44 as female, 2 as non-binary, and 8 did not disclose their gender. Regarding racial and ethnic identity, participants included 1 Native American, 45 Black, 31 White, 11 individuals identifying with other ethnicities, and 10 participants who chose not to report their ethnicity. Research and data collection protocol was conducted under approval from the Institutional Review Board of University of Pennsylvania and secondary analysis of the anonymized data was approved by the Institutional Review Board of University of Illinois at Urbana-Champaign. Participation was voluntary, and all students provided written assent along with parental consent.

Log data from WHIMC

This study used both log data and interview data to investigate interview reactivity. Log data were recorded from the Minecraft server and included students’ positional information sampled every two to three seconds (x, y, and z coordinates), records of tool use, and records of observational actions, each with corresponding timestamps.

DDCI from Quick Red Fox (QRF)

DDCI was conducted using Quick Red Fox (QRF), an open-source Android server–client application designed to notify interviewers when participants engage in behaviors of interest specified in advance by the research team [20]. QRF integrates modeling technologies and processes real-time data to identify key moments in learners’ experiences, referred to as interview triggers. Researchers can assign priorities to different triggers to detect rare or theoretically important behaviors while also distributing interview opportunities across students. When QRF detects a predefined behavioral pattern, the application notifies researchers, enabling them to decide whether and when to conduct an interview.

Trigger events and interview records are transmitted to a centralized dashboard, collecting logs of interview activity and allowing the research team to monitor interview activity in real time. The dashboard provides information such as trigger identifiers, timestamps, student and interviewer IDs, interviewer decisions to conduct interviews, and whether a second student contributed to an interview (e.g., when a neighboring student chimed in during an interview). Interviews are typically brief (generally under five minutes) to minimize disruption to students’ ongoing activities while capturing highly contextualized and ecologically valid accounts of their in-situ experiences. A total of four interviewers participate across the four camps (at least two per camp).

All interviews followed established DDCI interview principles and strategies [20], which emphasize maintaining respectful, non-evaluative relationships with students and eliciting the learners’ perspectives on their ongoing activities and experiences. Ocumpaugh et al. [20] describe this as the Big Sister Approach (BSA)--which they explicitly contrast with the Orwellian concept of Big Brother. Interviewers make it clear to the students before DDCI begins that interviewers are not there to enforce student behavior, and their interview approach is designed to reflect that as well. To that end, the BSA permits encouragement, but does not permit instruction or policing of off-task behaviors. An interviewer might offer a worried student gentle confirmation about whether the strategy the student wants to test might work, but students who need instruction should be referred to a teacher or peer during a DDCI.

For several reasons, DDCI blends a mix of questions about the learner and about the learning context. In part, questions about the learner allow the interviewer to build trust, however they are also part of an asset-based approach [21] where the interviewer is searching for knowledge that could be relevant to the learning context, even if the student has not yet made those connections. Although interviewers have access to the DDCI trigger information, they do not immediately ask the student about those constructs. The reason for this is two fold, since (1) it could make students feel surveilled, which could heighten self-presentation effects and (2) it limits the degree to which an interview can determine whether that construct and its associated experiences were salient to the students' game-play experience.

DDCI tends to start with questions like “how are you doing?” and “what strategies are you trying?” as they are open-ended enough to allow students to steer the topics in the initial part of each interview. As the interview progresses, questions may get more probing, and if the constructs associated with the trigger do not emerge, the interviewer may try to introduce them (e.g., asking if the student has deliberately visited any locations when determining how they came across a point of interest in the game). Throughout this process, however, the overarching goal is to encourage students to discuss their current states of metacognition, motivation, and engagement. If the constructs associated with the trigger (e.g., a particular part of the WHIMC game) are clearly not part of that metacognitive process, that is typically useful data for a DDCI study like this one, where the underlying goal is to understand student interest development. Therefore, in this study, interviewers did not always actively try to re-engage on topics directly associated with the trigger.

Pre-Interest Survey

The goal of QRF interviews in this research is to understand students' interest development through WHIMC. To account for individual differences, initial levels of individual interest were included in the analysis. Individual interest was measured using a survey adapted from Boeder et al. [2], which assesses context-independent interest in science across five subscales: Information Seeking, Motivation to Reengage, Persistence, Self-Regulation, and Value. Example items include “I want to know more about science” and “I seek out opportunities to engage in science.”

Each subscale comprised four to five items rated on a 7-point Likert scale, ranging from Do Not Agree At All to Very Strongly Agree. The survey was administered on the first day of the camp, prior to students’ engagement with the WHIMC activities.

Analysis

Among the log and interview data, only data from five What-If worlds (Lunar Crater, No Moon, Colder Sun, Tilted Earth, Two Moons) were included. Other activities (e.g., the Mars habitat activity and exoplanet exploration) were excluded due to differences in task designs.

The primary goal of the analysis was to examine whether interviews were associated with changes in students’ behaviors during and after the interview. To do so, we compared interviewed and non-interviewed students while controlling for pre-interview behavior. This approach helps disentangle interview effects from other influences such as observer effects or general time-related trends (e.g., declining activity as gameplay progresses) by considering both interview and non-interview cases. To enable direct comparison within a common temporal framework, we constructed a dummy interview time frame which aligned behavioral phases across interviewed and non-interviewed students. For each non-interviewed student in a given world, we determined their dummy interview start and end time by randomly selecting one of the interviewed students in the same world and camp and assigned that student’s interview start and end times to the non-interviewed student. In brief, we have a time framework composed of two parts : pre-interview (from world play start to interview start) and during/post-interview (from interview start to world play end).

This analysis focused on three core behaviors in WHIMC: tool use, observation, and movement. These behaviors represent central learning processes in the environment. Students move through the virtual world to locate relevant information, use scientific tools to explore phenomena, and make observations to document and reflect on their findings.

The number of tool use and observation behaviors were extracted directly from server logs without additional feature transformation. In addition to the count of observation and tool uses, we computed binary variables indicating whether the behavior occurred at least once within a given time window. This decision was motivated by the highly right-skewed count distributions for these behaviors, in which most observations were zero or near zero and a small number exhibited very high counts. Under such conditions, modeling only raw counts can be unstable and disproportionately influenced by rare extreme values. Accordingly, the analysis focused on not just frequency but also behavioral occurrence.

Movement-related features were engineered from positional log data from both phase (pre-interview and during/post-interview) and included mean distance traveled, mean speed, speed standard deviation, mean absolute acceleration, acceleration standard deviation, mean direction-change speed, direction-change speed standard deviation, and the mean of direction changes within three angular ranges (45–90°, 90–150°, and 150–180°).

Player position logs were first grouped per each user, camp, and world. All movement features were derived from two-dimensional spatial coordinates (x, z). To focus spatial navigation, only x (east-west) and z (north-south) coordinate was used while y axis (vertical) was excluded [33]. For each consecutive pair of position samples, stepwise displacement was computed as the Euclidean distance:

distancet=(Δxt)2+(Δzt)2\text{distanc}\text{e}_{\text{t}} = \sqrt{\left( \Delta x_{t} \right)^{2} + \left( \Delta z_{t} \right)^{2}} where Δxt=xtxt1\Delta x_{t} = x_{t} - x_{t - 1} and Δzt=ztzt1\Delta z_{t} = z_{t} - z_{t - 1}.

Instantaneous speed was then calculated as:

speedt=distancetΔtt\text{spee}\text{d}_{\text{t}} = \frac{\text{distanc}\text{e}_{\text{t}}}{\Delta t_{t}}

Acceleration was derived as the temporal derivative of speed:

accelt=ΔspeedtΔt{accel}_{t} = \frac{\Delta{speed}_{t}}{\Delta t}

Movement direction was computed at each step using the arctangent of the displacement vector:

headingt=atan2(Δzt,Δxt)\text{headin}\text{g}_{\text{t}} = atan2\left( \Delta z_{t},\Delta x_{t} \right), yielding angles in radians within (-π,π\pi,\ \pi].

Changes in heading between consecutive steps were calculated and wrapped to the principal angle range to account for circularity. The absolute turning angle (in degrees) was then defined as:

turn_degt=|headingtheadingt-1|×180π\text{turn\_de}\text{g}_{\text{t}} = \left| \text{headin}\text{g}_{\text{t}} - \text{headin}\text{g}_{\text{t-1}} \right| \times \frac{180}{\pi}

Angular turn velocity was computed as the rate of heading change per unit time:

ang_velt=ΔheadingtΔtt{\text{ang}\text{\_}\text{vel}}_{t} = \frac{\Delta\text{heading}_{t}}{\Delta t_{t}}

To characterize directional changes, turning events were categorized into three angular bins based on absolute turning angle: 45°–90°, 90°–150°, and 150°–180°

The data had a nested and longitudinal structure: each student participated in one of four camps, explored multiple worlds within the camp, and could be interviewed multiple times. To account for this dependency structure, we used generalized linear mixed-effects models (GLMMs) rather than standard regression. Models were fitted in R (version 4.5.2) using the glmmTMB package (version 1.1.13).

Depending on the distributional characteristics of features, we used different GLMM families. First, time-dependent features were fit with negative binomial GLMMs using during-post interview time as offset: the number of observations, and the number of tool use. Second, binary features are fit with binomial GLMMs: binary observation and binary tool use. Lastly, continuous features, which already accounted for time as part of the feature calculations, are modeled with normal distribution: mean distance, all speed features (mean speed and speed standard deviation), all acceleration features (mean absolute acceleration and acceleration standard deviation), and all angular features (mean absolute angular velocity, standard deviation of angular velocity, and mean of direction changes within three angular ranges (45–90°, 90–150°, and 150–180°)). After fitting the models, statistical significance was determined using the Benjamini–Hochberg false discovery rate correction to control for multiple comparisons and reduce Type I error.

A total of three analyses were conducted. First, because interviews were triggered by data-based behavioral criteria, we examined whether students who were interviewed differed significantly from those who were not in their behaviors prior to the interview. Also, as the goal of the interviews were to gain insights into the students' STEM interest, their pre-interest might influence whether they are interviewed. We tested whether interview status significantly predicted students’ pre-interview behaviors and pre-interview interest. In these models, camp and world were included as random effects.

Second, we compared the during-and-post features of interviewed and non-interviewed students while controlling for pre-behavior, interview, their interaction, and pre interest. Because we hypothesized that pre-interview behavior might moderate the effect of the interview, all models were initially fitted with an interaction term between pre-interview behavior and interview status. When neither the interaction nor the main effect of interview status was significant, an additional reduced model excluding the interaction term was fitted to examine potential main effects. Across all models, camp and world were included as random effects. When model convergence issues occurred, the random effect of camp was removed and the model was refitted.

Table 1. Extracted Interview Keywords.

Behaviors

Keywords

Observation

observation

Tool use

tool, measure, airflow, radiation, altitude, atmosphere, radius, tectonic, temperature, gravity, tide, humidity, tilt, magnetic, rotation, oxygen, year, pressure, composition, scale

Lastly, when interview effects were detected, we further tested whether these effects were driven by an attention-cue mechanism. Specifically, we compared students whose interviews explicitly mentioned a target behavior (e.g., observation) with those whose interviews did not mention that behavior. We determined whether an interview contained an attention cue by matching keywords associated to the behavior to the interview transcripts (Table 1). We used only ‘observation’ for observation as other possible keywords (e.g. observe) can extract many unrelated interviews while, for tool

Table 2. Mean and Standard Deviation of Features Before Interview for both Interviewed and Non-interviewed Students.

 

Feature

Interviewed

Not interviewed

N

Pre

During/Post

N

Pre

During/post

Absolute acceleration (mean)

79

1.31 (2.38)

1.17 (2.86)

194

1.04 (1.15)

1.19 (2.20)

Acceleration (std)

77

8.38 (24.57)

7.83 (26.15)

192

5.26 (12.36)

6.51 (19.02)

Angular absolute velocity (mean)

82

0.54 (0.33)

0.45 (0.40)

203

0.64 (0.41)

0.61 (0.43)

Angular velocity (std)

79

0.80 (0.44)

0.69 (0.51)

198

0.94 (0.51)

0.91 (0.54)

Distance (mean)

87

3.16 (6.35)

4.93 (12.13)

211

2.37 (6.61)

1.88 (5.43)

Speed (mean)

82

2.88 (2.08)

2.43 (2.16)

204

2.28 (1.61)

2.19 (1.89)

Observation (sum)

87

1.48 (2.15)

1.25 (2.09)

212

1.38 (4.10)

0.54 (1.44)

Observation (occurrence)

87

0.52 (0.50)

0.46 (0.50)

212

0.39 (0.49)

0.22 (0.41)

Speed (std)

79

9.67 (18.56)

8.86 (19.09)

198

6.19 (10.94)

6.55 (14.29)

Tool (sum)

87

1.75 (2.67)

0.77 (1.43)

212

1.58 (3.31)

0.58 (1.45)

Tool (occurrence)

87

0.59 (0.50)

0.32 (0.47)

212

0.37 (0.48)

0.23 (0.42)

Turn 150 to 180 (mean)

83

6.74 (5.37)

6.67 (8.63)

205

7.41 (6.03)

8.00 (6.65)

Turn 45 to 90 (mean)

83

9.45 (10.05)

7.09 (6.42)

205

11.42 (11.83)

9.57 (7.93)

Turn 90 to 150 (mean)

83

15.18 (12.03)

13.54 (13.00)

205

20.01 (15.29)

19.27 (15.44)

Table 3. General Linear Mixed Methods Results (Coefficients, Standard Errors, and P-values) for Models Predicting Each of the 14 Features.

Behavior

N

Pre-behavior

Interview

Pre-behavior * Interview

Pre-interest

Observation (sum)

276

β = 0.07 (0.05)

p = 0.154

β = 0.53 (0.33)

p = 0.108

β = -0.03 (0.12)

p = 0.836

β = 0.34 (0.09)

p = <.001

Tool (sum)

276

β = 0.12 (0.05)

p = 0.022

β = 0.32 (0.38)

p = 0.403

β = -0.12 (0.14)

p = 0.380

β = 0.38 (0.11)

p = <.001

Observation (occurrence)

276

β = 1.88 (0.40)

p = <.001

β = 1.30 (0.49)

p = 0.009

β = -0.51 (0.64)

p = 0.427

β = 0.16 (0.11)

p = 0.152

Tool (occurrence)

276

β = 1.09 (0.39)

p = 0.005

β = 1.41 (0.48)

p = 0.003

β = -1.65 (0.64)

p = 0.010

β = 0.33 (0.12)

p = 0.004

Absolute acceleration (mean)

231

β = 0.11 (0.19)

p = 0.550

β = -0.91 (0.43)

p = 0.035

β = 0.95 (0.28)

p = <.001

β = -0.10 (0.10)

p = 0.343

Acceleration (std)

229

β = -0.13 (0.19)

p = 0.515

β = -3.98 (3.05)

p = 0.191

β = 1.34 (0.29)

p = <.001

β = -0.05 (0.89)

p = 0.955

Angular absolute velocity (mean)

236

β = 0.61 (0.09)

p = <.001

β = -0.12 (0.09)

p = 0.172

β = 0.16 (0.13)

p = 0.224

β = -0.01 (0.01)

p = 0.500

Angular velocity (std)

231

β = 0.68 (0.11)

p = <.001

β = -0.04 (0.11)

p = 0.686

β = 0.02 (0.11)

p = 0.868

β = -0.01 (0.02)

p = 0.559

Distance (mean)

241

β = 0.30 (0.06)

p = <.001

β = -0.63 (0.69)

p = 0.361

β = 0.60 (0.09)

p = <.001

β = -0.02 (0.20)

p = 0.911

Speed (mean)

236

β = 0.54 (0.11)

p = <.001

β = 0.03 (0.45)

p = 0.946

β = -0.01 (0.14)

p = 0.935

β = 0.01 (0.08)

p = 0.921

Speed (std)

231

β = -0.00 (0.13)

p = 0.997

β = -1.81 (2.41)

p = 0.452

β = 0.43 (0.18)

p = 0.018

β = -0.10 (0.66)

p = 0.879

Turn 150 to 180

(frequency mean)

234

β = 0.39 (0.09)

p = <.001

β = -2.61 (1.45)

p = 0.072

β = 0.34 (0.15)

p = 0.026

β = -0.26 (0.29)

p = 0.359

Turn 45 to 90

(frequency mean)

234

β = 0.10 (0.05)

p = 0.038

β = -1.06 (1.33)

p = 0.425

β = 0.02 (0.09)

p = 0.853

β = -0.09 (0.31)

p = 0.760

Turn 90 to 150

(frequency mean)

234

β = 0.43 (0.08)

p = <.001

β = -1.49 (2.84)

p = 0.599

β = 0.02 (0.14)

p = 0.866

β = -0.37 (0.54)

p = 0.490

Note. β denotes fixed-effect estimates and values in parentheses are standard errors.

use, we included ‘tool’, ‘measure’, and other the name of available measurement tools. Because movement patterns cannot be reliably identified in interview transcripts, this analysis was limited to tool use and observation behaviors. After investigating direct attention cue effect, we refit the model in the second analysis (interview reactivity) while excluding interviews containing an attention cue for the target behaviors. This was done to see whether the interview effect is replicated with the condition. By reproducing interview reactivity, we can see whether direct attention cue effect can be divided with other potential mechanisms explaining the interview reactivity. Across all models, camp and world were included as random effects. When model convergence issues occurred, the random effect of camp was removed and the model was refitted.

RESULTS

Across the four camps, we collected 120 interviews from 60 students (mean: 2, std: 1.22) from the five target worlds. A total of 33 interviews were excluded because interview time logs (4 cases) or world playtime logs (9 cases) were unavailable; the interviews contained no meaningful content (7 cases), such as when students were chosen but they refused to participate or interviewers terminated the session within 50 seconds; or a student completed multiple interviews within the same world (20 cases), as repeated interviews within a world could distort the defined time frame. After screening the interviews, 87 interviews from 55 students remained (mean: 1.58, std: 0.85). Each camp had an average of 21.75 interviews (std: 8.06). The five worlds included in the analysis had an average of 17.4 interviews each (std: 2.30). The average interview time was 264.16 seconds (std : 117.90) while the average play time before an interview was 621.37 seconds (std : 372.32) and the average play time after an interview start was 653.14 (std : 530.16).

Behaviors Differences Before Interview

The first analysis examined whether students who were interviewed differed from those who were not interviewed in terms of their behavior or interest at a given time point. Table 2 provides descriptive statistics comparing behavior features for interviewed and non-interviewed students. GLMM models were used to test whether interview status significantly predicted students’ pre-interview behaviors and pre-interview interest. In these models, camp and world were included as random effects. The results indicated that interviewed students exhibited significantly higher occurrence of tool use (binary), whereas no other significant differences were observed between the interviewed and non-interviewed groups in either behavioral or interest measure.

Interview Reactivity

Second, we examined whether interview participation significantly predicted students’ behaviors after the interview. GLMM models were fit to compare during/post features while including pre-behavior, whether the student was interviewed or not, their interaction, and pre interest as fixed effect (Table 3).

Significant positive main effects of interviews were observed for several features. Specifically, interview participation significantly predicted the occurrence of observation, and occurrence of tool use. In these models, pre-interview behavior and pre-interview interest were also significant predictors, while occurrence of tool use had significant negative interaction effects. In case of observation occurrence, there is significant interview main effect without interaction effect (β = 1.30, SE = 0.49, p = 0.009). On the other hand, although interview was associated with an increased likelihood of tool use (β = 1.41, SE = 0.48, p = .003), this effect was significantly moderated by pre-interview tool use, as indicated by a negative Interview * Pre tool use interaction (β = −1.65, SE = 0.64, p = .010). This pattern suggests that interview reactivity may be most evident among students who had not used tools prior to the interview. As the main and interaction effects are pointed in different directions in tool use, interpreting the interview effect based solely on the main effect coefficient can be misleading. Therefore, we probed the interaction by examining the conditional pre–post behavioral association among interviewed students (i.e., the combined coefficient β_int + β_interaction). Simple effects analyses showed that interview participation did not significantly affect tool use among students who had already engaged in tool use before the interview (β = 0.24, SE = 0.42, p = .57). Together, these results highlight that students’ baseline behavioral status plays an important role in shaping interview reactivity.

Some movement features—including absolute acceleration mean, acceleration standard deviation, and mean distance—showed significant positive interaction effects with interview participation. However, although the corresponding main effects were not statistically significant, they were opposite in direction to the interaction terms. This pattern suggests that the direction and magnitude of interview-related effects may vary depending on students’ pre-interview behavioral levels. Overall, post-interview movement patterns remain systematically linked to baseline movement behaviors, implying that interview participation may partially shape subsequent movement dynamics.

For features that showed no interview or interaction effects in the full models, reduced models excluding the interaction term were fitted. However, no additional interview effects were detected for other features.

Direct Attention Cue Effect

Lastly, we examined whether the observed reactivity in the occurrence of observation and tool use was driven by direct attention cues during interviews. This was achieved by comparing interviews that explicitly mentioned the target behavior with those that did not. In terms of interviews mentioning observation and tool use, 31 interviews included the ‘observation’ keyword while 30 interviews mentioned tool use keywords (see table 1). We compared during/post features while including pre-behavior, interview mention, their interaction, and pre interest as fixed effect (Table 4). Camp and world were included as random effects across all models.

No significant direct attention cue effects or interaction effects with pre-behavior were observed for either observations or tool use. Since neither feature showed a main or an interaction effect in the full models, reduced models excluding the interaction term were fitted. However, no additional attention cue effects were detected for either observation or tool use.

After examining direct attention cue effects, we re-ran the analyses comparing the behaviors of students who were interviewed to that of students who did not get interviewed (as in section 4.2), this time excluding interviews which explicitly mentioned the target behavior (observation or tool use). These analyses were conducted to assess whether interview reactivity was still significant when only considering interviews that did not include direct attention cues. The results showed that the previously reported effect of interview on the occurrence of tool use and its interaction effect between interview and pre-behavior were fully replicated. In contrast, for the occurrence of observation, the previously identified main effect of interview participation was no longer observed.

Table 4. General Linear Mixed Methods Results (Coefficients, Standard Errors, and P-values) for Direct Attention Cue Models and Interview Effect Reproduction Models Predicting Each of Two Features: The Occurrence of Observation and Tool Use.

Behavior

N

Pre-

behv

Interview

Pre-behv * Interview

Pre-

interest

Direct attention cue effect

Observation (occurrence)

47

β = 1.42 (0.83)

p = 0.085

β = 1.48 (1.38)

p = 0.282

β = -2.28 (1.58)

p = 0.149

β = 0.14 (0.24)

p = 0.564

Tool

(occurrence)

47

β = -0.77 (1.18)

p = 0.513

β = -0.84 (1.00)

p = 0.401

β = 1.54 (1.32)

p = 0.242

β = 0.04 (0.26)

p = 0.879

Reproduction of interview effect without interview mention

Observation (occurrence)

257

β = 1.88 (0.42)

p = <.001

β = 1.11 (0.52)

p = 0.035

β = 0.16 (0.73)

p = 0.821

β = 0.15 (0.12)

p = 0.187

Tool (occurrence)

246

β = 1.10 (0.39)

p = 0.005

β = 1.43 (0.53)

p = 0.006

β = -2.14 (0.77)

p = 0.006

β = 0.33 (0.12)

p = 0.007

Note. β denotes fixed-effect estimates and values in parentheses are standard errors.

DISCUSSION AND CONCLUSION

This study examined whether and how data-driven classroom interview (DDCI) influences students’ behaviors in a game-based learning context, by comparing a range of in-game behaviors of students following their interviews to the behaviors of students who were not interviewed. Specifically, we investigated behavior differences related to the use of in game observations, scientific measurement tools and movement-related behaviors. The results indicate that DDCI does exert an influence on learners’ behaviors; however, this influence is selective rather than uniform across all behavioral features.

First, the occurrences of tool use and observation, two core inquiry behaviors in WHIMC, were positively influenced by interviews. This suggests that rather than distracting students from the learning activity, students may become more engaged in inquiry processes within the game environment following interviews. This finding aligns with the metacognitive reactivity described by Bosch et al. [3], highlighting DDCI’s potential to exert positive influences on students.

However, while students showed increased occurrences of tool use and observation following interviews, no significant interview effects and a few interaction effects were found for movement features. Reactivity in in-game observations and scientific measurement tool uses are particularly meaningful because they reflect intentional, goal-directed actions that learners may consciously engage in, in contrast to movement-related features that are more likely to emerge as byproducts of task engagement. This indicates that characteristics of each behavior such as intentionality or goal-direction may influence the impact of DDCI, as suggested by Double and Birney [4] that task characteristics influence how reactivity operates cognitive and metacognitive processes.

Even reactivity within a single behavior should be examined carefully, as it may involve complex underlying mechanisms. In our findings, interview reactivity predicted the occurrence of tool use and observation, but not the frequency of these behaviors. Even these occurrence-level effects were conditional: reactivity depended on students’ pre-interview behavior and was most pronounced among those who had not engaged in tool use prior to the interview. Together with other interaction effects for movement-related features in this study, these results suggest that reactivity is not uniform but varies according to individuals’ prior states [4]. This variability and complexity indicate that researchers should not conceptualize reactivity as a simple yes-or-no phenomenon, but rather as a multifaceted process shaped by the interplay of task characteristics and individual-level factors.

The mechanisms underlying the reactivity can be interpreted in multiple ways. In the present context, all students were aware of the presence of observers and interviewers; nevertheless, significant behavioral differences emerged between interviewed and non-interviewed students. This finding suggests that the effects observed cannot be explained solely by a general observer effect, but rather that the interview process itself is a contributing factor.

Research on verbalization reactivity supports this interpretation. Studies on think-aloud and verbal reporting have demonstrated that while spontaneous inner thought verbalization tends not to affect performance, higher-level verbalization—such as describing and explaining one’s current thoughts and actions—can be reactive [5, 7, 8]. In the present study, while our interviewers attempted to minimize disruption to learners’ ongoing activity by avoiding retrospective questions and instead asking students to describe what they were currently doing, responding to such prompts might still require learners to articulate and reflect on their ongoing behavior. Within Ericsson’s framework, this constitutes a form of reactive high-level verbalization, which may activate metacognitive processes by encouraging learners to reflect on their actions.

However, inconsistent with this interpretation, the present study didn’t identify attention cue effect described in the metacognitive reactivity framework proposed by Double and Birney [4]. No direct attention cue effect was revealed. Also, while interview reactivity was reproduced even without interviews mentioning the behaviors for the occurrence of tool use, this effect was not reproduced for the occurrence of observations. The success to reproduce reactivity of tool use occurrence suggests that the reactivity of tool use might be independent from attention cue effect. On the other hand, failure to reproduce reactivity of observation occurrence suggests either a broader form of attention cueing that is not captured by explicit behavioral mentions alone, or insufficient statistical power due to the small sample size to detect differences between cueing and non-cueing interviews.

These results may also reflect limitations in how attention cues were operationalized in this study. Rather than conducting a detailed analysis of interview content, we used a simple keyword matching approach to identify whether a given behavior was explicitly mentioned in an interview (e.g. mention of ‘observation’ or ‘tool’). It is possible that relevant behaviors were implicitly discussed without the use of specific keywords. For example, in the case of tool use, interviews may have referenced phenomena (e.g., windy area) or interpretations derived from tools without explicitly mentioning the tools themselves (e.g., ‘airflow’). Such instances were not systematically captured in the present analysis, which may have attenuated the attention cue effects. More fine-grained analyses of interview content are therefore needed to better understand the mechanisms of DDCI reactivity.

Given that both the present study and prior research [3] consistently indicated that students are likely to be reactive to DDCI, a critical methodological question concerns how such reactivity should be addressed. While traditional methodological perspectives have emphasized minimizing or eliminating reactivity [15], more recent work has advocated for a reactivity-transparent approach, in which reactivity is explicitly examined, documented, and analytically incorporated rather than treated solely as a source of bias [17, 34]. As DDCI enables the collection of ecologically valid data within authentic learning contexts, understanding how DDCI-induced reactivity influences research findings is essential to fully utilize DDCI’s power.

Researchers need to clarify what aspects of learners are influenced, how they are influenced, and when such influences occur. Beyond identifying the existence of reactivity, it is also critical to evaluate its implications for specific research questions. The same reactivity effect may have different consequences depending on the research objective. For example, when the goal is to model interest based on behavioral log, changes of their actual behavior may not substantially distort the results, and this can be tested by validating the model in another context. In contrast, when the research objective is explanatory—such as examining relationships among behaviors —asymmetric reactivity effects across behaviors may bias estimated relationships and lead to misleading conclusions.

Researchers should also consider how to mitigate or account for reactivity effects analytically. If reactivity primarily alters the overall occurrence of behaviors symmetrically across behavioral categories and individuals, normalized analytic approaches may help alleviate its impact [16]. For example, Epistemic Network Analysis (ENA), which emphasizes the relative weights of connections rather than raw behavioral frequencies, can reduce sensitivity to differences in behavioral volume across units. However, such approaches are less effective if reactivity alters the way behaviors are externally manifested, rather than merely their frequency.

Methods for the detailed analysis of students’ behavioral log data, such as the one developed by the educational data mining (EDM) community, can contribute to revealing more fine-grained measurement about reactivity and suggesting good ways to address it analytically. By leveraging the richness of behavioral log data and feature engineering techniques, researchers can move beyond treating reactivity as a uniform source and instead examine how it manifests selectively across different behavioral dimensions. In the present study, the use of multiple engineered features enabled us to identify asymmetric reactivity patterns in detail. More broadly, methodological approaches developed in the EDM community may provide analytic strategies which account for the influence of students’ reactivity to DDCI on the interpretation of student behavior patterns.

Although this study is among the relatively few to empirically examine reactivity associated with DDCI, it remains limited in scope, focusing on a restricted set of behaviors within a specific learning context. This limitation highlights the need for future research examining DDCI reactivity across diverse contexts and outcome measures, including behavioral, affective, and metacognitive indicators. Additionally, variations in reactivity across demographic groups should be considered, as certain groups may perceive greater levels of surveillance from interviewers. As such evidence accumulates, the role, strengths, and limitations of DDCI as a research instrument can be more clearly articulated. A similar process has unfolded in the development of think-aloud protocols. Since their introduction over three decades ago, think-aloud methods have been subject to sustained debate regarding their identity and reactivity. Through extensive empirical use, meta-analyses, and methodological debate, researchers have developed refined protocols aimed at minimizing reactivity and clarifying appropriate use cases [5–8, 28]. These findings played a central role in differentiating the method from other methods and in establishing empirically grounded guidelines for managing reactivity.

Similarly, DDCI requires sustained empirical investigation and cumulative synthesis to clarify its identity, strengths, and limitations relative to other research instruments. Beyond investigating how the interviews contribute to students’ reactivity during DDCI, future work should also examine how other components of the DDCI method, such as variations in interview length, interview triggers design or the interview prioritization algorithm, may impact student reactivity. Through such efforts, DDCI can be more rigorously positioned as a research tool, with its methodological affordances and constraints clearly understood.

ACKNOWLEDGMENTS

This study was supported by the National Science Foundation (NSF; DRL-2301172). Any conclusions expressed in this material do not necessarily reflect the views of the NSF. We would also like to thank the entire WHIMC project team for their support throughout this project.

Acknowledgement of generative ai use

This study used generative AI to assist with data analysis programming and draft revision. No study data were directly uploaded to AI systems, and all AI-generated content was carefully reviewed by the first author.

REFERENCES

  1. Baker, R.S., Hutt, S., Bosch, N., Ocumpaugh, J., Biswas, G., Paquette, L., Andres, J.M.A., Nasiar, N. and Munshi, A. 2024. Detector-driven classroom interviewing: focusing qualitative researcher time by selecting cases in situ. Educational technology research and development. 72, 5 (Oct. 2024), 2841–2863. https://doi.org/10.1007/s11423-023-10324-y.
  2. Boeder, J.D., Postlewaite, E.L., Renninger, K.A. and Hidi, S.E. 2021. Construction and validation of the Interest Development Scale. Motivation Science. 7, 1 (Mar. 2021), 68–82. https://doi.org/10.1037/mot0000204.
  3. Bosch, N., Zhang, Y., Paquette, L., Baker, R., Ocumpaugh, J. and Biswas, G. 2021. Students’ Verbalized Metacognition During Computerized Learning. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama Japan, May 2021), 1–12.
  4. Double, K.S. and Birney, D.P. 2019. Reactivity to Measures of Metacognition. Frontiers in Psychology. 10, (Dec. 2019). https://doi.org/10.3389/fpsyg.2019.02755.
  5. Ericsson, A. 2003. Valid and non-reactive verbalization of thoughts during performance of tasks towards a solution to the central problems of introspection as a source of scientific data. Journal of Consciousness Studies. 10, 9–10 (2003), 1–18.
  6. Ericsson, K.A. and Fox, M.C. 2011. Thinking aloud is not a form of introspection but a qualitatively different methodology: Reply to Schooler (2011). Psychological Bulletin. 137, 2 (2011), 351–354. https://doi.org/10.1037/a0022388.
  7. Ericsson, K.A. and Simon, H.A. 1980. Verbal reports as data. Psychological review. 87, 3 (1980), 215.
  8. Fox, M.C., Ericsson, K.A. and Best, R. 2011. Do procedures for verbal reporting of thinking have to be reactive? A meta-analysis and recommendations for best reporting methods. Psychological bulletin. 137, 2 (2011), 316.
  9. Gadbury, M. and Chad Lane, H. 2023. A Bayesian Analysis of Adolescent STEM Interest Using Minecraft. Artificial Intelligence in Education. Posters and Late Breaking Results, Workshops and Tutorials, Industry and Innovation Tracks, Practitioners, Doctoral Consortium and Blue Sky. N. Wang, G. Rebolledo-Mendez, V. Dimitrova, N. Matsuda, and O.C. Santos, eds. Springer Nature Switzerland. 384–389.
  10. Gieles, E., Vogelsmeier, L., Jongerling, J., Smeets, T. and Dejonckheere, E. 2025. Participation Effects in Ecological Momentary Assessment Research: A Taxonomy and Call to Action. PsyArXiv, 4 (Nov. 2025). https://doi.org/doi:10.31234/osf.io/cj8ut_v2.
  11. Greene, J.C. 2008. Is Mixed Methods Social Inquiry a Distinctive Methodology? Journal of Mixed Methods Research. 2, 1 (Jan. 2008), 7–22. https://doi.org/10.1177/1558689807309969.
  12. Koriat, A., Nussinson, R., Bless, H. and Shaked, N. 2013. Information-based and experience-based metacognitive judgments: Evidence from subjective confidence. Handbook of metamemory and memory. Psychology Press. 117–135.
  13. Kvale, S. and Brinkmann, S. 2009. Interviews: Learning the craft of qualitative research interviewing. sage.
  14. Lane, H.C., Gadbury, M., Ginger, J., Yi, S., Comins, N., Henhapl, J. and Rivera-Rogers, A. 2022. Triggering STEM interest with Minecraft in a hybrid summer camp. Technology, Mind, and Behavior. 3, 4 (2022).
  15. Lietz, C.A. and Zayas, L.E. 2010. Evaluating qualitative research for social work practitioners. Advances in Social work. 11, 2 (2010), 188–202.
  16. Liu, X., Gurung, A., Baker, R.S. and Barany, A. 2024. Understanding the Impact of Observer Effects on Student Affect. Advances in Quantitative Ethnography. Y.J. Kim and Z. Swiecki, eds. Springer Nature Switzerland. 79–94.
  17. Maxwell, J.A. 2013. Qualitative research design: An interactive approach: An interactive approach. sage.
  18. McCarney, R., Warner, J., Iliffe, S., Van Haselen, R., Griffin, M. and Fisher, P. 2007. The Hawthorne Effect: a randomised, controlled trial. BMC Medical Research Methodology. 7, 1 (Dec. 2007), 30. https://doi.org/10.1186/1471-2288-7-30.
  19. McCarthy, D.E., Minami, H., Yeh, V.M. and Bold, K.W. 2015. An Experimental Investigation of Reactivity to Ecological Momentary Assessment Frequency among Adults Trying to Quit Smoking. Addiction (Abingdon, England). 110, 10 (Oct. 2015), 1549–1560. https://doi.org/10.1111/add.12996.
  20. Ocumpaugh, J. et al. 2025. The Quick Red Fox gets the best Data Driven Classroom Interviews: A manual for an interview app and its associated methodology. arXiv.
  21. Ocumpaugh, J., Roscoe, R.D., Baker, R.S., Hutt, S. and Aguilar, S.J. 2024. Toward Asset-based Instruction and Assessment in Artificial Intelligence in Education. International Journal of Artificial Intelligence in Education. 34, 4 (Dec. 2024), 1559–1598. https://doi.org/10.1007/s40593-023-00382-x.
  22. Patton, M.Q. 2002. Qualitative research and evaluation methods 3rd. ed. Sage publications.
  23. Perera, K. 2020. The interview as an opportunity for participant reflexivity. Qualitative Research. 20, 2 (Apr. 2020), 143–159. https://doi.org/10.1177/1468794119830539.
  24. Procter, I. and Padfield, M. 1998. The effect of the interview on the interviewee. International Journal of Social Research Methodology. 1, 2 (Jan. 1998), 123–136. https://doi.org/10.1080/13645579.1998.10846868.
  25. Ravitch, S.M. and Carl, N.M. 2019. Qualitative research: Bridging the conceptual, theoretical, and methodological. Sage publications.
  26. Rosenthal, R. and Jacobson, L. 1966. Teachers’ Expectancies: Determinants of Pupils’ IQ Gains. Psychological Reports. 19, 1 (Aug. 1966), 115–118. https://doi.org/10.2466/pr0.1966.19.1.115.
  27. Saretsky, G. 1972. The OEO PC experiment and the John Henry effect. The Phi Delta Kappan. 53, 9 (1972), 579–581.
  28. Schooler, J.W. 2002. Verbalization produces a transfer inappropriate processing shift. Applied Cognitive Psychology. 16, 8 (Dec. 2002), 989–997. https://doi.org/10.1002/acp.930.
  29. Siemens, G. and Baker, R.S.J.D. 2012. Learning analytics and educational data mining: towards communication and collaboration. Proceedings of the 2nd International Conference on Learning Analytics and Knowledge (Vancouver British Columbia Canada, Apr. 2012), 252–254.
  30. Sporrong, S.K., Kalleberg, B.G., Mathiesen, L., Andersson, Y., Rognan, S.E. and Svensberg, K. 2022. Understanding and addressing the observer effect in observation studies. Contemporary research methods in pharmacy and health services. Elsevier. 261–270.
  31. Wengraf, T. 2001. Qualitative research interviewing: Biographic narrative and semi-structured methods. (2001).
  32. Winne, P.H. 2020. Construct and consequential validity for learning analytics based on trace data. Computers in Human Behavior. 112, (Nov. 2020), 106457. https://doi.org/10.1016/j.chb.2020.106457.
  33. Zhou, Y. and Paquette, L. 2024. Investigating Student Interest in a Minecraft Game-Based Learning Environment: A Changepoint Detection Analysis. Proceedings of the 17th International Conference on Educational Data Mining, EDM 2024 (Atlanta United States, July. 2024). https://doi.org/10.5281/ZENODO.12729844.
  34. Zahle, J. 2023. Reactivity and good data in qualitative data collection. European Journal for Philosophy of Science. 13, 1 (Mar. 2023), 10. https://doi.org/10.1007/s13194-023-00514-z.

© 2026 Copyright is held by the author(s). This work is distributed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0) license.