ABSTRACT
This study examines the stability and instructional context sensitivity of keystroke writing process features across a full academic year of middle school writing. Using 8,576 in-class writing sessions collected from two U.S. middle schools, we analyze whether writing-process indicators remain consistent within students over time and how they vary across instructional contexts. We define stability as within-student consistency of writing features across tasks and sensitivity as systematic variation associated with grade level and school. A broad set of features capturing fluency, pausing, revision, and engagement was extracted from students’ longitudinal writing portfolios. We quantified stability using intraclass correlation coefficients and examined contextual sensitivity using linear mixed-effects models with random student-level intercepts. Results show that most features exhibit low stability within School x Grade cohorts, indicating substantial within-student variability across tasks. Only a small subset of low-level timing features related to word-level text production demonstrated moderate stability, whereas many revision- and engagement-related features varied systematically across instructional contexts. These findings clarify the measurement properties of keystroke-derived writing indicators and inform their appropriate use in longitudinal educational data mining applications.
Keywords
1. INTRODUCTION
Writing is a complex cognitive activity that involves coordinating planning, text production, and revision [13]. In educational settings, understanding how students engage in these processes is essential for assessing writing development and supporting instruction. However, traditional writing research has largely relied on final texts or observational methods, which provide limited insight into how writing unfolds over time [23]. Recent advances in keystroke logging technologies enable fine-grained capture of writers’ real-time behaviors by recording keystrokes, pauses, and edits during composition [2, 22]. This data offers direct access to writing processes that are invisible in finished texts. These processes however, have become increasingly important in educational data mining (EDM) for applications such as learner modeling, performance prediction, and process-based feedback [7, 14].
Prior work shows that keystroke-derived features are educationally meaningful [37]. More proficient writers tend to produce text more fluently and pause strategically, whereas less proficient writers pause more frequently and revise differently [1, 33]. At the same time, writing processes are highly context-sensitive. Features such as total writing time, revision frequency, and output vary substantially across tasks, genres, and instructional conditions, while only a subset of low-level timing measures exhibit moderate stability across occasions [8, 32].
These findings highlight a fundamental measurement challenge for longitudinal writing analytics: some process features reflect relatively persistent student-level tendencies, whereas others primarily capture contextual adaptation [8]. This distinction matters even within a single instructional setting. As students encounter different tasks, genres, and instructional demands across an academic year, observed changes in their writing process features may reflect shifts in task context rather than genuine developmental growth. Without first establishing which features are stable within students across tasks and which vary systematically with context, longitudinal analyses risk conflating developmental change with task-specific effects, leading to misleading inferences about writing growth [2]. This motivates the present study’s focus on measurement before modeling, characterizing the stability and contextual sensitivity of keystroke-derived features as a foundation for future longitudinal work.
This challenge is particularly salient for educational applications that aim to model writing development over time or deploy process-based analytics in classrooms [7, 14]. Before modeling growth or generating feedback, it is necessary to establish whether keystroke-derived features are interpretable across writing samples [2, 32]. Accordingly, this study adopts a measurement-focused perspective that precedes growth modeling. Rather than modeling change directly, we examine the stability and contextual sensitivity of writing process indicators in authentic classroom settings.
We address this goal through two research questions:
- RQ1: To what extent do keystroke-derived writing process indicators exhibit within-student consistency across tasks in students’ longitudinal writing portfolios?
- RQ2: To what extent does variability in these indicators reflect systematic sensitivity to instructional context, specifically grade level, incoming academic proficiency, and school membership?
To answer these questions, we analyze longitudinal keystroke data collected from everyday classroom writing across an entire school year in two middle schools. By characterizing which process features exhibit within-cohort stability and which vary systematically with context, this work provides a principled foundation for longitudinal learner modeling and process-aware educational data mining in real-world instructional environments [2, 36].
2. BACKGROUND
2.1 Writing process features
Research on writing processes has consistently shown that keystroke-derived features capture meaningful aspects of how students compose text, including productivity, fluency, pausing, and revision behavior [22, 33]. Across populations, more proficient writers tend to write more fluently, produce longer bursts of text, and pause less frequently, reflecting more efficient translation of ideas into language [15, 18]. Fluency-related measures, in particular, have been repeatedly linked to writing quality and performance in both classroom and assessment contexts [5].
Pausing behavior provides additional insight into cognitive processing during composition. Skilled writers typically exhibit shorter pauses at character, word, and sentence boundaries. This is consistent with more automatic lexical access and smoother coordination between planning and production [26, 35]. Burst length, the amount of text produced between pauses, has emerged as a robust indicator of writing fluency, with longer bursts associated with higher-quality writing across multiple studies [2, 5].
Revision behavior presents a more complex picture. Prior work suggests that stronger writers engage in more strategic, higher-level revisions that reorganize or clarify meaning, while less skilled writers tend to make frequent surface-level edits during drafting [11, 34]. Although revision features are informative, their interpretation often depends strongly on task demands and instructional context [9]. Together, this body of work demonstrates that keystroke-derived features capture meaningful variation in writing processes and motivate applications in learner modeling, performance prediction, and process-based feedback within educational data mining [7, 16].
2.2 Stability of writing process features
A central methodological question in writing process research concerns whether keystroke-derived features represent stable characteristics of writers or fluctuate substantially across tasks. Early work emphasized that many process measures reflect multiple interacting cognitive and situational factors and cautioned against assuming direct mappings to stable traits [16, 28]. However subsequent empirical studies have demonstrated that stability varies widely across features.
In assessment contexts for example, Deane and Zhang showed that only a limited subset of keystroke features exhibits moderate test– retest reliability across essay prompts, while many others display low consistency [37]. Crucially, these stable features were also the ones that generalized across tasks in predictive models, whereas highly variable features failed to support robust inference about writing ability [10]. This finding underscores that stability is not merely a psychometric concern but a prerequisite for longitudinal interpretation and modeling.
Longitudinal evidence from text-based writing research reinforces the importance of this distinction between stability and change. Using mixed-effects models over a large learner corpus, Lei, Wen, and Yang showed that development in syntactic complexity unfolds gradually and is characterized by substantial within-writer variability over time [21]. Their findings emphasize that longitudinal inference depends on features that are sufficiently stable to support growth modeling, while still allowing for uneven, context-dependent trajectories of change.
Related results have been reported in adult and second-language writing. Choi and Deane found that despite substantial within-writer variability, certain low-level timing features remained relatively stable across tasks and time points and were meaningfully related to writing quality [6]. Across studies, a consistent pattern emerges: low-level fluency measures (e.g., typing speed, within-word pause duration) tend to behave as writer-specific tendencies, whereas higher-level features related to planning, revision, or sentence-level pausing are more sensitive to task characteristics [1, 37].
2.3 Contextual sensitivity and task effects
Alongside work on stability, a substantial literature demonstrates that writing processes are highly sensitive to task and context. Writing behavior varies with prompt complexity, genre, topic familiarity, time constraints, and instructional conditions [4, 17]. Experimental studies manipulating task difficulty show that output-related features and revision frequency increase substantially for more demanding tasks, while core timing features such as average inter-key interval remain comparatively stable [1].
At the same time, longitudinal writing development is inherently context-sensitive. A study demonstrated that trajectories of syntactic complexity growth vary systematically with instructional and contextual factors, even when modeled at scale [21]. This finding highlights a key challenge for longitudinal writing research: apparent change may reflect shifts in task demands or learning contexts rather than underlying development, underscoring the need to interpret longitudinal patterns in relation to both stability and context.
Genre effects further highlight contextual sensitivity. Argumentative writing typically elicits longer pauses and more revisions than narrative writing, reflecting increased planning and evaluation demands [4, 11]. In assessment settings, researchers have cautioned that process measures observed for one prompt may not generalize to another, limiting their interpretability across occasions [10, 36]. For longitudinal analyses, this presents a key challenge: observed changes in process features may reflect contextual variation rather than genuine development.
2.4 Measurement & modeling
Interest in writing process data has grown across research on learner modeling and instructional support. Keystroke-derived features have been incorporated into predictive models of writing quality and shown to add explanatory power beyond static text features [6, 7]. Process features have also been explored as signals for real-time feedback and formative analytics in writing support systems [22, 29].
Broader efforts in human-computer interaction have also examined the design of intelligent writing support systems. A study proposed a comprehensive design space for writing assistants organized around user, task, technology, interaction, and ecosystem dimensions, highlighting writing efficiency and proficiency as key user-facing considerations [20]. The measurement properties of keystroke-derived features examined in the present study are directly relevant to such systems, as reliable process signals are a prerequisite for effective real-time writing support.
However, a recurring concern in this line of work is the validity of process features across contexts. If the meaning of a feature varies substantially across genres or task conditions, its use in general-purpose models becomes problematic [1, 36]. Modeling applications therefore require a principled understanding of which features are robust across contexts and which require contextualized interpretation.
Prior research establishes that writing process features capture skill-related patterns, vary widely in stability, and are strongly influenced by task context. What remains underexplored is how these features behave across broader instructional contexts, such as grade level and school, within naturalistic longitudinal classroom data. This study addresses that gap by systematically examining the stability and contextual sensitivity of keystroke-derived writing process features in everyday classroom writing, providing a measurement foundation for longitudinal modeling and process-aware analysis.
3. DATASET
3.1 Data Collection
This study uses longitudinal keystroke data collected via Writing Observer, a research platform integrated into Google Docs and deployed in two U.S. public middle schools. Writing Observer is a suite of tools designed to support classroom orchestration in writing-to-learn tasks. The platform unobtrusively captures real-time writing data at the keystroke level, recording students’ editing behaviors, reading activity, and document interactions from start to finish of each writing session. It is also designed to integrate into existing learning management systems to track class enrollments, assignments, and evaluations, enabling monitoring of writing practices within and across assignments. Since students contributed multiple writing sessions over an academic year, the resulting dataset supports analysis of within-student consistency and contextual variation in writing-process features across authentic classroom tasks.
All data collection was conducted in accordance with institutional research ethics requirements. The study operated under IRB approval, with students participating under appropriate consent and assent procedures. All records were pseudonymized prior to analysis, and no personally identifiable information is reported.
3.2 Participants
The study participants included 423 middle school students in Grades 7 and 8 from two U.S. public schools who participated during the 2023–2024 academic year. One school was a large urban middle school in Utah, and the other was a suburban middle school in Pennsylvania. Together, these schools provided a diverse instructional context while remaining comparable in grade span and curricular structure.
The writing activities were part of regular classroom instruction, primarily in English Language Arts (ELA) courses, with additional writing in subjects such as science and social studies. Assignments included narrative, argumentative, and informative writing, as well as text-dependent analysis, and non-ELA writing such as lab reports and class notes. Session durations varied and were computed from the first to the last logged keystroke event within each document. Writing prompts were not standardized and varied naturally by grade, teacher, genre, and topic, reflecting authentic classroom practice rather than controlled assessment conditions.
We collected a total of 9,831 writing-session logs. After excluding documents with fewer than 50 words, the final corpus consisted of 8,576 sessions (3,641 from School 1 and 4,935 from School 2). All data were collected under appropriate institutional approvals, and logs were de-identified prior to analysis. Table 1 summarizes sample sizes by school, grade, and dataset. Figures 1 and 2 show the distribution for both schools.
Also, we collected the incoming proficiency scores of all the students from the 2022-2023 academic year which were binned into four categories (1 = Below Basic, 2 = Basic, 3 = Proficient, 4 = Advanced). Given the absence of Below-Basic students in School 2 and their low numbers in School 1, Below Basic group was excluded to maintain comparability across schools.
Gender distribution across cohorts was as follows: School 1 Grade 7 had 70 female and 40 male students; School 1 Grade 8 had 47 female and 51 male students; School 2 Grade 7 had 31 female and 30 male students; and School 2 Grade 8 had 75 female and 79 male students.
| School | Grade 7 | Grade 8
| ||
| Students | Docs | Students | Docs | |
| 1 | 110 | 1840 | 98 | 1801 |
| 2 | 61 | 989 | 154 | 3946 |


3.3 Feature Extraction
Each keystroke log consists of a timestamped sequence of writing events, including the character insertions, deletions, pauses, and navigation actions. Pauses were inferred from inter-keystroke intervals (IKIs), and revision behavior was identified through insertion–deletion patterns in the event stream [2, 22].
From these logs, we derived a set of keystroke-based process features describing how students generated, revised, and interacted with text. Following prior writing-process research (Section 2), features were organized into four behavioral dimensions: fluency, revision, complexity, and engagement [1, 8, 10, 36]. In total, ten frequency-based features (capturing how often actions occurred) and fourteen time-based features (capturing pause durations and temporal dynamics) were computed. We summarize the feature families:
- Fluency: Features capturing transcription efficiency and writing flow, including the number of inserts, in-word inserts, between-word inserts, before-word inserts, word-start inserts, and pause measures such as IKI between inserts, within-word IKI, between-word IKI, before-word IKI, and word-start IKI [1, 22].
- Revision: Features describing editing and text modification behavior, including before-word inserts, total edits (insertions + deletions), starting and continuing backspaces, edit duration, and pause measures associated with deletion activity (IKI starting backspace, IKI continuing backspace) [2, 10].
- Complexity: Features reflecting structural and syntactic elaboration during writing, including clause-break inserts, minor punctuation inserts (e.g., commas, parentheses), and associated pause measures at clause boundaries and minor punctuation (IKI at clause break, IKI at minor punctuation) [1, 4].
- Engagement: Features that capture non-typing interactions with the document, including overall menu time, reading time (inactive but focused time), navigation time (cursor, scrolling, arrow keys), and total active writing time [35].
Time-based features were log-transformed using \(\log (1 + x)\) to mitigate heavy-tailed pause distributions, and pauses longer than 120 seconds were truncated to exclude off-task intervals [1, 22]. For each applicable time-based feature, the mean, median, standard deviation, and variance were computed, yielding a total of 44 features used in analysis [2, 10]. Although a broader feature set was available, this study focuses on core indicators of fluency, revision, complexity, and engagement that are most relevant to examining longitudinal stability and contextual sensitivity [1, 8].
4. METHODS
Our analysis addresses two research questions using complementary statistical approaches. RQ1 examines the stability of keystroke-derived writing process features within cohorts using variance decomposition via intraclass correlation, while RQ2 examines sensitivity to contextual factors using linear mixed-effects modeling [12, 27].
4.1 Data Preparation and Cohort Definition
Writing process features were extracted from document-level writing logs, with multiple writing samples available per student. Because proficiency-level information was not available at the document level, student proficiency-level labels were obtained from a separate roster summary and merged onto document-level records using anonymized student identifiers.
Writing samples without an associated proficiency-level label were excluded from further analysis. Analyses were conducted separately for each school, and within each setting, data were further stratified by grade level. All results reported for RQ1, therefore, reflect within-cohort analyses, defined by a specific combination of school and grade. To ensure reliable estimation of student-level effects, analyses were restricted to students who contributed at least two writing sessions within a given cohort, and only features present and numerically valid within that cohort were included.
4.2 RQ1: Intra-Class Correlation Analysis for Feature Stability
RQ1 investigates the extent to which keystroke-derived writing process features exhibit stable, student-specific structure within cohorts, as opposed to variability across writing samples produced by the same student. Stability is quantified using the intraclass correlation coefficient (ICC), which decomposes total variance into between-student and within-student components [12, 24]. In the context of writing-process analysis, stability indicates whether a feature reflects a persistent student-level tendency rather than transient task-specific behavior.
For each writing process feature and for each school × grade cohort, we estimated the intraclass correlation coefficient (ICC) using a random-intercept formulation:
where \(y_{ij}\) denotes the value of a given writing-process feature for student \(i\) on writing task \(j\), \(\mu \) is the overall mean, \(u_i \sim \mathcal {N}(0, \sigma ^2_{\text {student}})\) represents a student-specific random intercept capturing stable between-student differences, and \(\varepsilon _{ij} \sim \mathcal {N}(0, \sigma ^2_{\text {within}})\) captures within-student variability across tasks, , where \(\sigma ^2_{\text {student}}\) and \(\sigma ^2_{\text {within}}\) denote the between-student and within-student variance components, respectively.
We computed the intraclass correlation coefficient (ICC) as:
which represents the proportion of total variance attributable to stable between-student differences. Under this random-intercept formulation, total variance is fully decomposed into between-student and within-student components, such that their sum represents the total variability in the feature. An ICC close to 1 indicates that a feature is highly stable within students across writing tasks, whereas an ICC near 0 indicates that most variability arises from task-to-task fluctuations within students.
For each feature, point estimates of the ICC were obtained using a random-intercept mixed-effects model estimated via restricted maximum likelihood (REML) [31]. When model estimation was unstable, point estimates were computed using an ICC estimator based on a one-way analysis of variance [30]. To quantify uncertainty, we computed ICCs and corresponding 95% confidence intervals for each feature using all available writing sessions. Following established conventions, ICC values were categorized as indicating poor (\(< 0.20\)), fair (\(0.20\)–\(0.40\)), moderate (\(0.40\)–\(0.60\)), or high (\(\ge 0.60\)) stability [19].
This analysis enables ranking writing-process features by their degree of longitudinal consistency and identifying which features function as persistent student-level signatures versus those dominated by task-specific variation. Establishing this distinction is a prerequisite for meaningful longitudinal modeling of writing development, as features that lack sufficient stability cannot be reliably interpreted as indicators of student growth.
4.3 RQ2: Modeling Contextual Sensitivity
RQ2 examines how writing-process features vary systematically with school context, grade level, and academic proficiency, while accounting for repeated observations from the same student [31, 3]. Specifically, RQ2 asks whether a portion of observed feature variance can be explained by known contextual factors rather than treated as undifferentiated task noise. Each feature was modeled using a linear mixed-effects framework with fixed effects for contextual variables and a random student-level intercept. Although writing-process features may exhibit complex or nonlinear patterns over time, the present analysis uses a linear mixed-effects framework because contextual predictors are modeled as categorical fixed effects, allowing for flexible, non-monotonic differences across groups without imposing a specific functional form.
For each writing-process feature, we fit the following model:
where \(y_{ij}\) is the feature value for student \(i\) on writing task \(j\), \(\beta _0\) is the intercept, and \(\beta _1\), \(\beta _2\), and \(\beta _3\) represent fixed effects of school, grade level, and incoming academic proficiency, respectively. The term \(b_i \sim \mathcal {N}(0, \sigma ^2_{\text {student}})\) captures student-specific deviations from the population mean, and \(\varepsilon _{ij} \sim \mathcal {N}(0, \sigma ^2_{\text {residual}})\) represents residual variability [31, 3].
We treated school, grade level, and incoming academic proficiency as categorical fixed effects to allow for non-linear and non-monotonic differences across groups [31]. Academic proficiency was operationalized using a categorical state assessment score with four ordered levels.
To identify school effects, we fit models across schools and grade levels, ensuring that the school predictor varied within each subset. We estimated all models using restricted maximum likelihood and included only students who had at least two writing sessions [31]. Subsets with insufficient numbers of students or insufficient variation in predictors were excluded.
For each fitted model, we report fixed-effect estimates, standard errors, z-statistics, and approximate two-sided p-values based on normal approximations [29]. We also report variance components for the student-level random intercept and the residual term to contextualize the relative contribution of between- and within-student variability.
5. RESULTS
5.1 RQ1: Feature Stability Within Cohorts
This section reports results for RQ1, which examines the stability of keystroke-derived writing process features within School × Grade cohorts. We quantify stability using intraclass correlation coefficients (ICCs), where higher values indicate greater consistency of a feature across multiple writing samples produced by the same student, relative to variability between students within the same cohort.
Across all cohorts, a substantial proportion of features yielded ICC values of 0 (or effectively 0), indicating no detectable between-student variance after accounting for within-student variability. Tables 2–5 report only features with non-zero ICC estimates; all omitted features had ICC = 0 and are therefore not interpretable as stable individual-level measures in these contexts. Overall, feature-level stability varied markedly across schools, grade levels, and feature types. The majority of features exhibited Poor stability, with only a small subset, primarily low-level insertion timing features, demonstrating Moderate or High stability, most consistently in School 2.
5.1.1 School 1 – Grade 7
In School 1, Grade 7, the writing process features demonstrated overwhelmingly low stability (Table 2). Of the 41 features examined, most produced ICC values of 0, indicating that variability was almost entirely within students across writing tasks rather than between students. Among the features with non-zero ICCs, only meanLogWordStartInsertTime reached Moderate stability (ICC = 0.448). meanLogBeforeWordInsertTime showed Fair stability (ICC = 0.288), while all remaining reported features, including measures of timing variability, backspacing behavior, and navigation time, exhibited Poor stability (ICC < 0.20).
These results suggest that, in this cohort, most writing-process features do not function as stable student-level characteristics and instead fluctuate substantially across writing tasks.
| Feature | ICC | Stability |
|---|---|---|
| meanLogWordStartInsertTime | 0.448 | Moderate |
| meanLogBeforeWordInsertTime | 0.288 | Fair |
| stdLogBetweenWordInsertTime | 0.184 | Poor |
| varLogContinuingBackspaceTime | 0.104 | Poor |
| varLogEditTime | 0.094 | Poor |
| varLogBeforeWordInsertTime | 0.082 | Poor |
| stdLogWordStartInsertTime | 0.074 | Poor |
| overallNavigationTime | 0.056 | Poor |
5.1.2 School 1 – Grade 8
In School 1, Grade 8, stability patterns remained limited, though a small number of features demonstrated modest increases relative to Grade 7 (Table 3). As in Grade 7, the majority of the features analyzed yielded ICCs of 0 and are not shown. Two within-word insertion timing features, meanLogInWordInsertTime (ICC = 0.494) and medianLogInWordInsertTime (ICC = 0.484), demonstrated Moderate stability, while medianLogStartingBackspaceTime reached Fair stability (ICC = 0.386). All other reported features exhibited Poor stability, and no features achieved High stability.
Overall, while some fine-grained timing features became more consistent in Grade 8, feature-level stability in School 1 remained sparse and highly selective.
| Feature | ICC | Stability |
|---|---|---|
| meanLogInWordInsertTime | 0.494 | Moderate |
| medianLogInWordInsertTime | 0.484 | Moderate |
| medianLogStartingBackspaceTime | 0.386 | Fair |
| meanLogClauseBreakInsertTime | 0.175 | Poor |
| varLogContinuingBackspaceTime | 0.144 | Poor |
5.1.3 School 2 – Grade 7
In contrast, School 2 Grade 7 exhibited markedly higher stability across several features (Table 4). Although many of the 41 features still yielded ICC values of 0, a larger subset demonstrated meaningful between-student consistency. Three features, medianLogInWordInsertTime (ICC = 0.681), meanLogInWordInsertTime (ICC = 0.679), and meanLogBetweenWordInsertTime (ICC = 0.620), reached High stability, indicating strong within-student consistency across writing sessions. Several additional features related to word-start insertion and starting backspace timing demonstrated Moderate stability (ICC \(\approx \) 0.50–0.59). Despite this, higher-order variability measures (e.g., variance-based timing features and punctuation timing variability) continued to show Poor stability or zero ICCs.
These results suggest that, in this context, micro-temporal production behaviors may reflect relatively stable individual writing signatures, while broader variability-based features do not.
| Feature | ICC | Stability |
|---|---|---|
| medianLogInWordInsertTime | 0.681 | High |
| meanLogInWordInsertTime | 0.679 | High |
| meanLogBetweenWordInsertTime | 0.620 | High |
| medianLogBetweenWordInsertTime | 0.594 | Moderate |
| medianLogWordStartInsertTime | 0.589 | Moderate |
| meanLogWordStartInsertTime | 0.567 | Moderate |
| medianLogStartingBackspaceTime | 0.504 | Moderate |
| meanLogStartingBackspaceTime | 0.462 | Moderate |
| meanLogBeforeWordInsertTime | 0.266 | Fair |
| stdLogMinorPunctInsertTime | 0.145 | Poor |
| varLogContinuingBackspaceTime | 0.142 | Poor |
| varLogBeforeWordInsertTime | 0.042 | Poor |
| stdLogClauseBreakInsertTime | 0.032 | Poor |
5.1.4 School 2 – Grade 8
In School 2, Grade 8, overall stability declined relative to Grade 7, though it remained higher than in School 1 (Table 5). As in other cohorts, many of the 41 features analyzed had ICCs of 0 and are omitted. Two within-word insertion timing features, meanLogInWordInsertTime (ICC = 0.589) and medianLogInWordInsertTime (ICC = 0.568), retained Moderate stability, while meanLogClauseBreakInsertTime demonstrated Fair stability (ICC = 0.315). All remaining reported features exhibited Poor stability, and no features reached High stability.
These findings indicate that even within the same school context, feature stability can attenuate with grade level and that stable features become increasingly restricted to specific low-level timing measures.
| Feature | ICC | Stability |
|---|---|---|
| meanLogInWordInsertTime | 0.589 | Moderate |
| medianLogInWordInsertTime | 0.568 | Moderate |
| meanLogClauseBreakInsertTime | 0.315 | Fair |
| stdLogWordStartInsertTime | 0.124 | Poor |
| stdLogContinuingBackspaceTime | 0.106 | Poor |
| stdLogClauseBreakInsertTime | 0.059 | Poor |
| stdLogStartingBackspaceTime | 0.054 | Poor |
| overallActiveWritingTime | 0.015 | Poor |
5.1.5 Summary of RQ1 Findings
Across all School × Grade cohorts, there are three key patterns that emerged:
- Zero ICCs are the norm: In every cohort, a substantial number of the 41 analyzed features yielded ICC values of 0, indicating that most writing-process features do not exhibit stable between-student differences.
- Stability is highly feature-specific: When stability
emerged, it was largely confined to within-word and word-initial insertion timing features, rather than variability based or aggregate measures. - Stability is context-dependent: School 2, particularly Grade 7, exhibited systematically higher stability than School 1, suggesting that instructional or institutional context may shape the consistency of writing processes.
Taken together, these results indicate that most keystroke-derived writing process features should not be treated as stable individual traits. Instead, stability appears rare, selective, and context-dependent, motivating the mixed-effects analyses in RQ2, which explicitly examine how contextual factors explain systematic variation in writing process behavior.
5.2 RQ2: Sensitivity to Contextual Factors
RQ2 examined whether variability in keystroke-derived writing process features can be systematically explained by contextual factors, specifically, school, grade level, and incoming academic proficiency, after accounting for repeated observations from the same student. To address this question, we fit a series of linear mixed-effects models, one per feature, including fixed effects for school, grade, and proficiency level, and a random intercept for student. All models were fit on the Full dataset, pooling Grades 7 and 8 across both schools.
5.2.1 Overall Patterns of Contextual Sensitivity
Across features, contextual sensitivity varied substantially by feature type. In general, incoming academic proficiency exhibited more consistent associations with writing process behavior than grade level and school membership. This pattern suggests that observed variation in writing processes reflects proficiency-related differences more than developmental or school-specific artifacts. Random intercept variances for students were consistently non-zero across models, confirming substantial between-student variability and justifying the use of mixed-effects modeling to account for repeated writing samples per student. Throughout the following subsections, fixed-effect coefficients are reported alongside back-transformed percentage differences (computed as \(e^{\hat {\beta }} - 1\)), which reflect multiplicative change in the original pause durations given that features were log-transformed prior to modeling.
5.2.2 Grade-Level Effects
Grade-level effects were observed for several timing-based features. In particular, Grade 8 students often exhibited shorter within-word insertion times than Grade 7 students. These effects were typically negative in direction, indicating increased efficiency or fluency with advancing grade level. However, grade effects were not universal across all features and were generally modest in magnitude, suggesting that developmental differences manifest more strongly in fine-grained timing behaviors than in global duration measures. Table 6 summarizes features with significant grade-level effects.
| Feature | \(\beta \) | Interpretation |
|---|---|---|
| Median In-word Pause | \(-0.03\) | 3% shorter |
| Mean In-word Pause | \(-0.03\) | 3% shorter |
| Std. In-word Pause | \(-0.02\) | 2% less varying |
| Median Clause-break Pause | \(-0.05\) | 5% shorter |
| Mean Clause-break Pause | \(-0.05\) | 5% shorter |
| Var. Clause-break Pause | \(-0.21\) | 19% less varying |
| Median Edit Pause | \(+0.02\) | 2% longer |
5.2.3 Proficiency Effects
Academic proficiency demonstrated clearer and more interpretable associations with writing process features than grade level. For multiple fine-grained timing features, the highest proficiency level (Level 4) was associated with systematically faster within-word insertion behaviors relative to the Basic proficiency group, while Level 3 effects were generally small and non-significant. These patterns indicate that keystroke-level timing measures capture meaningful differences in writing fluency linked to incoming academic proficiency. Importantly, proficiency effects were not uniformly incremental across adjacent levels. In several cases, only the highest proficiency group differed significantly from the reference group, indicating potential threshold effects rather than linear progression across proficiency bands. Table 7 summarizes features with significant proficiency effects.
| Feature | \(\beta \) | Interpretation |
|---|---|---|
| Median In-word Pause | \(-0.04\) | 4% shorter |
| Mean In-word Pause | \(-0.05\) | 5% shorter |
| Std. In-word Pause | \(-0.02\) | 2% less varying |
| Median Clause-break Pause | \(-0.08\) | 8% shorter |
| Mean Clause-break Pause | \(-0.08\) | 8% shorter |
| Var. Clause-break Pause | \(+0.19\) | 21% more varying |
| Mean Minor Punct. Pause | \(-0.080\) | 7.7% shorter |
5.2.4 School Effects
School membership explained relatively little variance in writing process features after controlling for grade and proficiency. Most school-related coefficients were small and non-significant, indicating that differences in writing processes were largely consistent across institutional contexts. This finding supports the generalizability of keystroke-based process measures across schools and suggests that observed behavioral differences are not driven by local instructional or institutional idiosyncrasies. Table 8 summarizes the two features with significant school effects.
| Feature | \(\beta \) | Interpretation |
|---|---|---|
| Std. In-word Pause | \(-0.02\) | 2% less varying |
| Mean Minor Punct. Pause | \(-0.09\) | 8.6% shorter |
5.2.5 Summary of RQ2 Findings
Taken together, the RQ2 results demonstrate that while many writing process features exhibit substantial within-student variability, a subset shows systematic sensitivity to meaningful contextual factors, particularly incoming academic proficiency and, to a lesser extent, grade level. Fine-grained timing features appear especially informative for distinguishing writing behaviors associated with higher proficiency, whereas global duration measures reflect engagement-related differences. These findings complement the stability analyses in RQ1 by showing that writing process features may be individually variable yet still contextually interpretable when modeled appropriately.
6. DISCUSSION
This study examined keystroke-derived writing-process features from two complementary perspectives: their stability across students and tasks (RQ1) and their systematic sensitivity to contextual factors such as grade level, school, and incoming academic proficiency (RQ2). Taken together, the findings reveal a nuanced picture in which most writing process features are highly variable at the individual level, yet a subset remains meaningfully interpretable when contextualized appropriately.
6.1 Limited Stability of Writing Process Features
The RQ1 analyses demonstrate that zero or near-zero ICCs were the norm across all School × Grade cohorts. For a large proportion of the 41 features examined, between-student variance was effectively absent once within-student variability across writing tasks was accounted for. This finding challenges the common assumption, implicit in many learning analytics applications, that keystroke-derived features can be treated as stable, student-specific traits.
When stability did emerge, it was highly feature-specific, largely confined to low-level temporal production measures, such as within-word and word-initial insertion times. These features plausibly reflect motor–cognitive coordination processes that are less sensitive to task content and genre, making them more likely to persist across writing contexts. In contrast, higher-order variability measures, revision timing dispersion, and aggregate duration features consistently exhibited poor or zero stability, suggesting that these measures are strongly task-dependent rather than student-specific.
Importantly, stability was also context-dependent. School 2, particularly Grade 7, exhibited substantially higher stability than School 1, indicating that instructional context, curriculum alignment, or writing task design may shape the consistency with which students express their writing processes. This variability cautions against generalizing stability findings from a single institutional setting and underscores the importance of reporting stability estimates alongside proposed writing-process indicators. This finding does not contradict the RQ2 result showing no significant school effect. The two analyses ask different questions about the data: RQ2 examines whether students in different schools write differently on average, while RQ1 examines whether individual students write consistently across tasks. A school context can produce similar average writing behaviors across institutions while still shaping how consistently individual students express those behaviors from task to task.
6.2 Reconciling Instability With Contextual Sensitivity
At first glance, the widespread instability observed in RQ1 may appear to undermine the usefulness of keystroke-derived writing process features. However, the RQ2 results clarify that lack of individual-level stability does not imply lack of interpretability. While many features fluctuate substantially across students within tasks, a subset shows systematic sensitivity to meaningful contextual factors, particularly incoming academic proficiency and grade level.
Fine-grained timing features, especially within-word and word-initial insertion measures, showed consistent associations with both grade and proficiency. Higher-grade and higher proficiency students tended to exhibit faster insertion behaviors, aligning with theoretical accounts of increased writing fluency and automatization with development and skill acquisition. These findings suggest that although such features may not function as stable individual signatures, they nonetheless capture reliable group-level differences tied to developmental and proficiency-related factors.
By contrast, macro-level duration features, such as overall active writing time, followed a different pattern. These features showed limited grade sensitivity but were associated with proficiency in ways that suggest differences in engagement or persistence rather than efficiency. Notably, higher-proficiency students spent more time actively writing, indicating that longer writing durations should not be interpreted simplistically as inefficiency, but rather as potentially reflecting sustained effort or deeper engagement.
This finding resonates with and extends prior work on revision behavior in writing support systems. A study demonstrated that revision patterns captured from writing tool logs are meaningful and vary systematically across learners, including differences tied to gender and feedback condition [25]. Our results complement this by showing that revision features are among the least stable across tasks at the individual level, suggesting that longitudinal pipelines relying on revision indicators should account for substantial within-student variability rather than treating revision behavior as a fixed learner characteristic.
6.3 Implications for Writing Analytics and Educational Dashboards
Together, these findings have important implications for the design and interpretation of writing analytics systems. First, they suggest that keystroke features should not be universally treated as stable learner traits, especially when used for longitudinal profiling or individual diagnosis. Most features are better understood as context-sensitive behavioral indicators that respond to task demands, instructional conditions, and developmental factors.
Second, the results highlight the value of mixed-effects modeling frameworks for research on the writing process. By explicitly modeling both within-student variability and contextual predictors, such approaches allow researchers and practitioners to extract meaningful patterns from features that would otherwise appear noisy or unreliable at the individual level.
Finally, the differential behavior of fine-grained timing features versus aggregate duration measures suggests that feature families serve distinct interpretive roles. Low-level timing features may be well-suited for examining fluency and automatization, while duration-based features may provide insight into engagement, persistence, or strategic effort. Analytics systems and dashboards should therefore avoid aggregating heterogeneous features into undifferentiated indices and instead support feature-specific interpretation.
7. CONCLUSIONS
This paper examined the measurement properties of keystroke-derived writing process features in authentic classroom settings, focusing on their within-cohort stability and sensitivity to contextual factors. Using a large corpus of middle school writing sessions, we showed that most writing process features exhibit low or zero stability within School × Grade cohorts, reflecting substantial within-student variability across writing tasks. Only a small subset of fine-grained timing features related to word-level text production demonstrated moderate stability, and high stability was rare and strongly context-dependent.
In contrast, many writing process features varied considerably across writing sessions. Mixed-effects analyses revealed that, for a limited set of features, this variability aligns systematically with incoming academic proficiency and grade level, rather than reflecting undifferentiated noise. In particular, within-word insertion timing features showed consistent associations with proficiency, with higher-proficiency students exhibiting faster text production. Grade-level effects were present but modest, suggesting that writing process development during middle school unfolds gradually and that much observed variability reflects situational adaptation rather than rapid developmental change. After controlling for grade and proficiency, school-level effects were small and largely non-significant, supporting the generalizability of keystroke-based process measures across institutional contexts.
Taken together, these findings clarify an important measurement distinction: keystroke-derived writing-process features do not uniformly serve as stable student-level indicators. Some features provide relatively consistent signals within cohorts, while others primarily capture context-responsive behavior. This distinction is critical for longitudinal research, as features dominated by within-student variability cannot be interpreted as reliable indicators of enduring writing characteristics when analyzed in isolation.
The primary contribution of this study is methodological. Rather than modeling writing growth directly, we establish which classes of process features are sufficiently stable to support longitudinal interpretation and which require explicit contextual modeling. This groundwork enables more principled feature selection for future growth modeling, evaluation, and intervention studies in educational data mining.
From a practical perspective, the results suggest complementary roles for different feature types in educational tools. Features with higher within-cohort stability may support longitudinal monitoring of foundational writing processes, whereas context-sensitive features may serve as indicators of task-specific difficulty, instructional demands, or adaptive strategy use. Interpreting these signals jointly allows educators and researchers to distinguish enduring writing patterns from situational challenges.
While future work is needed to link process features to writing quality, extend analyses across longer time scales, and replicate findings in broader populations, this study provides a focused and principled analysis of writing-process measurement in real classroom contexts. Overall, the results demonstrate that keystroke-derived writing data contain both reliable structure and meaningful variability. Advancing longitudinal writing analytics will depend on leveraging the former while modeling the latter with appropriate attention to context.
7.1 Limitations and Future Work
While this study provides a systematic characterization of the stability and contextual sensitivity of keystroke-derived writing process features, several limitations should be acknowledged, along with directions for future research.
7.1.1 Absence of Writing Outcome Measures
A primary limitation of this study is the absence of direct measures of writing quality or learning outcomes. Although we identify features that exhibit greater stability within School × Grade cohorts (RQ1) and others that vary systematically with contextual factors such as grade level and incoming academic proficiency (RQ2), we do not evaluate whether these patterns correspond to improved writing performance or desirable writing behaviors. Stability alone does not imply educational value, and context sensitivity does not necessarily indicate instructional relevance.
Without outcome measures such as essay scores, rubric-based assessments, or teacher feedback, it is not possible to determine whether stable features reflect productive writing habits, emergent fluency, or merely idiosyncratic behaviors. Future work will integrate outcome measures to examine how stable and context-sensitive process features relate to writing quality and development. In particular, longitudinal growth models linking changes in writing process features to changes in writing performance will be critical for validating the educational significance of the process indicators identified in this study.
7.1.2 Limited Temporal Scope
The data analyzed in this study span approximately one academic year, which constrains conclusions about long-term stability and developmental change. Features that appear stable over this time frame may still evolve over longer periods, particularly across major instructional transitions or shifts in curricular emphasis. As a result, stability within a single school year should not be interpreted as permanence or trait-like invariance.
Longer-term longitudinal studies, especially those tracking students across multiple grade levels or school transitions, would help determine how writing process features change over extended time scales. Such studies could reveal developmental trajectories, periods of consolidation, or phases of heightened sensitivity to instruction that are not observable within a single academic year.
7.1.3 Unmodeled Task-Level Variation
Writing data were collected in authentic classroom settings, resulting in substantial variation in task topics, genres, instructional goals, and scaffolding. While this ecological validity strengthens the relevance of the findings, it introduces task-level variability that cannot be fully disentangled from student-level variability. In the absence of repeated or standardized prompts, it remains difficult to separate task-specific effects from within-student fluctuation, particularly for features that exhibited low or zero stability in RQ1.
Additionally, temporal distribution data for writing sessions across the academic year was not recorded, limiting our ability to assess whether uneven clustering of sessions could have introduced bias in ICC estimation. Task-level metadata, including genre classification and prompt type, was also not fully available at the time of analysis, constraining interpretation of school-level effects and task-driven variability. Future work will incorporate richer task annotation as labeling efforts are completed.
A related limitation concerns the inclusion of non-ELA writing tasks, such as lab reports and class notes, alongside ELA assignments. While this reflects the authentic range of classroom writing students engage in, it introduces additional heterogeneity that may affect certain feature families disproportionately. For instance, complexity features designed to capture syntactic elaboration may be less interpretable for highly formulaic or note-taking tasks.
Future research could adopt hybrid study designs that combine naturalistic classroom writing with repeated benchmark tasks, and examine whether stability patterns hold when analyses are restricted to ELA writing. Such designs would allow more direct estimation of test–retest reliability while preserving ecological validity. Additionally, richer annotation of task characteristics, such as genre, prompt complexity, revision expectations, and instructional supports, would enable more precise modeling of context-sensitive writing behaviors across subject areas.
7.1.4 Future Research Directions
Building on the findings of this study, several directions for future research are planned. First, we will integrate writing quality outcomes to model how stable and context-sensitive process features relate to writing development over time. Second, stability estimates from RQ1 can be leveraged to design process-aware educational tools, such as dashboards that emphasize meaningful deviations from a student’s typical writing behavior rather than raw feature values.
Finally, future work will investigate the sources of stability and variability through richer contextual modeling and qualitative inquiry, and will extend the feature set to include linguistic, syntactic, and higher-level process indicators. Together, these efforts aim to advance a more nuanced understanding of writing processes as both persistent and adaptive, supporting longitudinal modeling and actionable educational data mining applications.
8. REFERENCES
- V. M. Baaijen and D. Galbraith. Discovery through writing: Relationships with writing processes and text quality. Cognition and Instruction, 36(3):199–223, 2018.
- V. M. Baaijen, D. Galbraith, and K. De Glopper. Keystroke analysis: Reflections on procedures and measures. Written Communication, 29(3):246–277, 2012.
- D. J. Barr, R. Levy, C. Scheepers, and H. J. Tily. Random effects structure for confirmatory hypothesis testing: Keep it maximal. Journal of memory and language, 68(3):255–278, 2013.
- C. Beauvais, T. Olive, and J.-M. Passerault. Why are some texts good and others not? relationship between text quality and management of the writing processes. Journal of Educational Psychology, 103(2):415–428, 2011.
- N. A. Chenoweth and J. R. Hayes. Fluency in writing: Generating text in l1 and l2. Written communication, 18(1):80–98, 2001.
- I. Choi and P. Deane. Evaluating writing process features in an adult efl writing assessment context: A keystroke logging study. Language Assessment Quarterly, 18(2):107–132, 2021.
- R. Conijn, C. Cook, M. van Zaanen, and L. Van Waes. Early prediction of writing quality using keystroke logging. International Journal of Artificial Intelligence in Education, 32(4):835–866, 2022.
- R. Conijn, J. Roeser, and M. Van Zaanen. Understanding the keystroke log: The effect of writing task on keystroke features. Reading and Writing, 32(9):2353–2374, 2019.
- P. Deane. Writing assessment and cognition. ETS Research Report Series, 2011(1):1–60, 2011.
- P. Deane and M. Zhang. Exploring the feasibility of using writing process features to assess text production skills. ETS Research Report Series, 2015(2):1–16, 2015.
- L. Faigley and S. Witte. Analyzing revision. College Composition & Communication, 32(4):400–414, 1981.
- J. L. Fleiss and J. Cohen. The equivalence of weighted kappa and the intraclass correlation coefficient as measures of reliability. Educational and psychological measurement, 33(3):613–619, 1973.
- L. Flower and J. R. Hayes. A cognitive process theory of writing. College Composition & Communication, 32(4):365–387, 1981.
- H. Guo, P. D. Deane, P. W. van Rijn, M. Zhang, and R. E. Bennett. Modeling basic writing processes from keystroke logs. Journal of Educational Measurement, 55(2):194–216, 2018.
- J. R. Hayes. Modeling and remodeling writing. Written communication, 29(3):369–388, 2012.
- J. R. Hayes and L. S. Flower. Identifying the organization of writing processes. In L. W. Gregg and E. R. Steinberg, editors, Cognitive processes in writing, pages 3–30. Routledge, New York, NY, 2016.
- R. T. Kellogg. Training writing skills: A cognitive developmental perspective. Journal of writing research, 1(1):1–26, 2008.
- R. T. Kellogg. A model of working memory in writing. In C. M. Levy and S. Ransdell, editors, The science of writing, pages 57–71. Routledge, New York, NY, 2013.
- T. K. Koo and M. Y. Li. A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of chiropractic medicine, 15(2):155–163, 2016.
- M. Lee, K. I. Gero, J. J. Y. Chung, S. B. Shum, V. Raheja, H. Shen, S. Venugopalan, T. Wambsganss, D. Zhou, E. A. Alghamdi, T. August, A. Bhat, M. Z. Choksi, S. Dutta, J. L. Guo, M. N. Hoque, Y. Kim, S. Knight, S. P. Neshaei, A. Shibani, D. Shrivastava, L. Shroff, A. Sergeyuk, J. Stark, S. Sterman, S. Wang, A. Bosselut, D. Buschek, J. C. Chang, S. Chen, M. Kreminski, J. Park, R. Pea, E. H. R. Rho, Z. Shen, and P. Siangliulue. A design space for intelligent and interactive writing assistants. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI ’24, pages 1054:1–1054:35, New York, NY, USA, 2024. Association for Computing Machinery.
- L. Lei, J. Wen, and X. Yang. A large-scale longitudinal study of syntactic complexity development in efl writing: A mixed-effects model approach. Journal of Second Language Writing, 59:100962, 2023. Article 100962.
- M. Leijten and L. Van Waes. Keystroke logging in writing research: Using inputlog to analyze and visualize writing processes. Written Communication, 30(3):358–392, 2013.
- C. M. Levy and S. Ransdell. The science of writing: Theories, methods, individual differences and applications. Routledge, 2013.
- K. O. McGraw and S. P. Wong. Forming inferences about some intraclass correlation coefficients. Psychological methods, 1(1):30–46, 1996.
- L. Mouchel, T. Wambsganss, P. Mejia-Domenzain, and T. Käser. Understanding revision behavior in adaptive writing support systems for education. 2023.
- T. Olive and R. T. Kellogg. Concurrent activation of high-and low-level production processes in written composition. Memory & Cognition, 30(4):594–600, 2002.
- J. C. Pinheiro and D. M. Bates. Mixed-effects models in S and S-PLUS. Springer, 2000.
- G. Rijlaarsdam and H. Van den Bergh. Writing process theory. In C. A. MacArthur, S. Graham, and J. Fitzgerald, editors, Handbook of writing research, pages 41–53. Guilford Press, New York, NY, 2006.
- A. Shibani, S. Knight, and S. Buckingham Shum. Contextualizable learning analytics design: A generic model and writing analytics evaluations. In Proceedings of the 9th international conference on learning analytics & knowledge, pages 210–219, 2019.
- P. E. Shrout and J. L. Fleiss. Intraclass correlations: uses in assessing rater reliability. Psychological bulletin, 86(2):420–428, 1979.
- T. A. Snijders and R. Bosker. Multilevel analysis: An introduction to basic and advanced multilevel modeling. 2011.
- M. A. Tabari, X. Lu, and Y. Wang. The effects of task complexity on lexical complexity in l2 writing: An exploratory study. System, 114:103021, 2023.
- L. Van Waes and M. Leijten. Fluency in writing: A multidimensional perspective on writing fluency applied to l1 and l2. Computers and Composition, 38:79–95, 2015.
- L. Van Waes and P. J. Schellens. Writing profiles: The effect of the writing mode on pausing and revision patterns of experienced writers. Journal of pragmatics, 35(6):829–853, 2003.
- Å. Wengelin. Examining pauses in writing: Theory, methods and empirical data. In K. P. H. Sullivan and E. Lindgren, editors, Computer key-stroke logging and writing, pages 107–130. Elsevier, Oxford, 2006.
- M. Zhang and P. Deane. Process features in writing: Internal structure and incremental value over product features. ETS Research Report Series, 2015(2):1–12, 2015.
- M. Zhu, M. Zhang, and P. Deane. Analysis of keystroke sequences in writing logs. ETS Research Report Series, 2019(1):1–16, 2019.
© 2026 Copyright is held by the author(s). This work is distributed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0) license.