ABSTRACT
This study investigates the cognitive and behavioral characteristics of Chinese digital writing by modeling Pinyin-mediated keystroke processes. Prior keystroke research has largely focused on alphabetic languages, where keystrokes map directly onto text, limiting transferability to Chinese input methods that involve phonological entry followed by candidate selection. To address this gap, we propose a Dual-layer Rhythm Model (DLRM) that separates physical keystroke execution from cognitive-level selection and commitment processes. Using fine-grained keystroke logs from 34 graduate students completing both a transcription task and an argumentative writing task, we extracted process features related to input rhythm, candidate-selection latency, pause structure, and revision behavior. Within-subject comparisons showed that argumentative writing produced significantly greater rhythm variability, longer selection latencies, and higher revision rates than transcription, indicating increased cognitive load. Correlation analyses further revealed that cognitive-layer pause indicators, especially within-character pauses and post-commit pauses, were significant negative predictors of argument quality, whereas physical-layer speed measures were not. These findings demonstrate that cognitively grounded timing features derived from Pinyin input provide meaningful indicators of writing proficiency and support the development of process-aware models for Chinese writing assessment and intelligent writing support systems.
Keywords
INTRODUCTION
Writing involves complex cognitive processes such as planning, organizing, expressing ideas, and revising[6]. These processes require the coordinated operation of higher-order functions, including working memory, attentional control, and knowledge retrieval[6, 25]. Therefore, inferring writers’ cognitive abilities solely from the final text provides an incomplete and potentially biased view, process-level evidence is also needed[8]. A deeper understanding of the cognitive mechanisms underlying the writing process is essential for strengthening the validity of writing assessment, improving writing instruction, and developing automated, process-sensitive interventions[10].
Traditional process-tracing methods such as think-aloud protocols are limited by verbalization costs, incomplete introspective access, and poor scalability[13, 20].
Keystroke logs provide a window into cognitive processes that are not visible in the final written product [20]. They capture fine-grained behaviors such as typing speed, pauses, revisions, and bursts of text production[7, 20]. Moreover, the high temporal resolution, non-intrusive nature, and ecological validity of keystroke data make it well suited for large-scale studies of writing processes and for developing automated, process-aware writing assessment systems[8, 20].
However, research on keystroke-log analysis for Chinese writing remains far less developed than in English[11]. Chinese writers typically rely on Pinyin input methods as a proxy for text production[34]. Most Chinese input methods involve a two-stage process: writers first enter Pinyin letters that represent the intended Chinese characters, and then identify and select the target character from a list of candidates[21]. This workflow introduces additional cognitive steps, making analytical approaches developed for alphabetic-language keystroke data difficult to transfer to Chinese. Consequently, advancing keystroke-based methods for studying Chinese writing requires process-analytic techniques that explicitly model the Pinyin-mediated input stream[23].
Empirical research on Chinese writing with Pinyin input remains limited. Existing studies have primarily examined how Pinyin input affects lexical processing, while paying less attention to how Pinyin-mediated input unfolds during authentic writing tasks[6]. Systematic analyses that link input-process features to writing quality are especially scarce. Against this backdrop, the present study proposes a method for analyzing keystroke logs of Pinyin-mediated Chinese input, characterizes behavioral features of the input process, and investigates their underlying cognitive mechanisms. We further model the relationship between input processes and writing quality. Specifically, we address the following research questions:
RQ1: What are the key characteristics of Pinyin-mediated Chinese input, and how should these features be interpreted?
RQ2: How does input during writing differ from input during simple typing, and what cognitive activities might these differences reflect?
RQ3: Which input features are associated with writing proficiency, and what are their plausible cognitive interpretations?
RELATED WORK
Chinese Pinyin Keystroke Analysis
Keystroke-log research in English writing has established relatively mature theoretical and methodological foundations[20]. In English, keystroke events map relatively transparently onto the language production process. Accordingly, English keystroke-log studies have focused on physical-level properties of letter input[9], focusing on physical-level properties such as IKI, pause frequency, revision rate, and burst length [7, 20].
However, entering Pinyin letters externalizes only phonological encoding; the intended Chinese characters are not committed to the text until the writer identifies and selects the target orthographic form from a candidate list (see Figure 1)[16, 26, 28, 38]. This adds a cognitive step from sound-based encoding to character-form selection that does not exist in alphabetic keyboard input.

Consider intra-word pauses as an example. In Chinese keystroke data, there is no direct counterpart to the English operational definition[16, 29, 30]. Some studies have imported English criteria and defined intra-word pauses as pauses when the cursor is positioned between two Chinese characters within a “word” in the text box[36]. Theoretically, however, such pauses are difficult to interpret as spelling retrieval or orthographic encoding, making the validity of these analyses questionable. More importantly, these definitions ignore pauses at the Pinyin-entry stage and pauses during candidate selection[35].
In Pinyin-mediated input, physical keystrokes represent only the first stage of language production. The cognitive translation from Pinyin to Chinese characters has its own temporal dynamics and cognitive demands[15]. The interval between completing Pinyin entry and confirming a character (i.e., selection latency) is a Pinyin-specific temporal indicator that reflects the cost of orthographic recognition and decision making. Such cognitively meaningful indicators may contain critical information about Chinese writing processes, yet prior Chinese keystroke analyses have not systematically accounted for them[32]. Moreover, relying on cursor position in the text box introduces an additional challenge: accurate Chinese word segmentation, which is itself non-trivial and can propagate error into feature extraction. Together, these issues create substantial technical and theoretical challenges for keystroke-based research on Chinese writing.
More importantly, prior work rarely uses designs that disentangle the cognitive load induced by transcription from the cognitive load of the writing task itself[18, 19]. Keystroke logs reflect an intertwining of input and composition processes: inter-keystroke intervals are shaped by typing proficiency, but they also vary with writers’ ongoing planning, formulation, and revision[20, 33]. In Pinyin-mediated Chinese input, transcription itself involves multi-stage cognitive processing and substantial individual differences. Without controlling for, or directly measuring, the cognitive demands of the input system, it is difficult to estimate the cognitive characteristics of the writing process accurately[14, 37].
Working-Memory Model of Writing
Baddeley’s multicomponent model of working memory provides a useful framework for explaining the cognitive mechanisms involved in writing[3]. The model comprises three subsystems: the phonological loop, the visuospatial sketchpad, and the central executive[4, 5]. These subsystems interact to enable written composition.
Pinyin-mediated input places qualitatively different demands on working memory than handwriting[24]. During phonological encoding, Pinyin input requires writers to activate phonological representations of the intended words and maintain them in the phonological loop[27]. In contrast, during handwriting, many writers rely more on visuospatial memory or directly map orthographic forms to motor actions[17]. During candidate selection, Pinyin input requires maintaining a visual representation of the candidate list in the visuospatial sketchpad while the central executive supports feature comparison and choice decisions. During strategy management, the central executive must monitor the input method’s predictions, evaluate the efficiency of the current strategy, and switch strategies when needed[12]. Consequently, especially when the process is not fully automatized, Pinyin input may impose higher working-memory load and reduce the resources available for other components of writing.
Skill acquisition theory posits that complex skills shift from controlled to automatic processing[1, 2, 22] When Pinyin spelling and candidate selection are highly automatized, transcription costs are minimized, freeing cognitive resources for content development[25] . In contrast, non-automatized input competes with composing processes, disrupting idea flow[19, 31]
Building on the cognitive theories reviewed above, we propose the DLRM (Dual-layer Rhythm Model) of Pinyin input as a theoretical framework for analyzing Chinese digital writing processes. The model decomposes Pinyin-mediated input into two relatively independent temporal dimensions: a physical layer and a cognitive layer. The physical layer refers to inter-keystroke intervals between Pinyin letters, which reflect the automatization of Pinyin encoding and the fluency of finger-motor execution. These temporal patterns are primarily shaped by input proficiency; once a sufficient level of skill is reached, Pinyin entry becomes a highly automatized motor program whose rhythm is relatively stable across writing tasks and less sensitive to fluctuations in task-related cognitive load.
In contrast, the cognitive layer captures the interval from completing Pinyin entry to confirming the target Chinese character, as well as intervals between successive character confirmations. These time points mark key cognitive transitions from phonological activation to orthographic retrieval, while also carrying signals of higher-level composing processes such as lexical access, syntactic assembly, and semantic integration. Consequently, cognitive-layer timing is expected to be highly sensitive to writing-task demands. The relationship between the two layers reflects writers’ resource allocation strategies: skilled writers can maintain physical-layer fluency while allocating sufficient time at the cognitive layer to support high-quality decisions.
Based on this framework, we propose the following hypotheses:
H1: Compared with a transcription task, an essay-writing task will impose higher cognitive load, reflected in (a) greater variability in physical-layer rhythm, (b) longer cognitive-layer processing time, and (c) more frequent revisions and strategic input behaviors.
H2: Cognitive-layer indicators (e.g., selection latency, inter-character confirmation intervals, and pause patterns during selection) will predict argumentative writing quality more strongly than physical-layer indicators.
METHOD
Participants and Experimental Design
We used a within-subject experimental design to characterize Pinyin-mediated Chinese input behaviors associated with idea generation under different writing-task conditions. Participants were 39 native Chinese-speaking graduate students, all of whom had extensive experience using Pinyin input methods for computer-based writing. The study was conducted in a controlled, in-person setting. Keystroke data were recorded using the logging module integrated into the CLASS platform.
The experiment included two writing tasks designed to induce different levels of cognitive load during Chinese Pinyin input (see appendix A and B).
The first task was a transcription task, in which participants were asked to accurately copy a provided Chinese philosophical passage within five minutes. The text discussed the nature of logic and knowledge and was approximately 173 characters long. This task served as a low–cognitive-load baseline condition, primarily reflecting transcription and input-rhythm characteristics under minimal idea-generation demands.
The second task was a timed argumentative writing task. Participants were given a structured argument prompt adapted from a standardized analytical writing assessment and were asked to compose a full written response within 30 minutes. The prompt required participants to analyze the assumptions underlying a recommendation and evaluate its logical validity. This task was designed to elicit high-level cognitive activities, including planning, reasoning, and argument construction, under authentic composition conditions.
Data Collection and Preprocessing
Before feature computation, we cleaned and standardized the raw keystroke logs. We excluded one participant with extremely low task completion and four participants who did not produce valid Pinyin-mediated input, yielding a final analytic sample of 34 participants. We then harmonized potential logging differences caused by input-methods by retaining only events directly related to text production. Finally, we reconstructed each participant’s continuous input stream by ordering events chronologically within each task. After cleaning, Task 1 contained 21,975 valid keystroke records and Task 2 contained 99,554 records.
Modeling Pinyin Input Behaviors
Because Pinyin-mediated Chinese input does not directly produce Chinese characters but proceeds through Pinyin letter entry followed by candidate selection, we explicitly modeled the Pinyin-mediated input process. Using the keystroke logs, we reconstructed each participant’s stream of Pinyin letter entry by tracking insertions and deletions at the end of the Pinyin buffer and assigning a timestamp to each letter event. We then computed inter-key intervals (IKIs) at the Pinyin-letter level and extracted summary statistics to characterize input rhythm and its stability. In addition, we quantified deletion behavior as an indicator of online monitoring during transcription.
In Pinyin-mediated input, the commitment of a Chinese character to the final text is treated as a key boundary of text production. End-of-text generation typically reflects the immediate externalization of ideas: writers must concurrently translate concepts into linguistic form, retrieve lexical items, and perform syntactic encoding. Its temporal properties therefore provide a more direct indicator of cognitive load during content generation.
Accordingly, we focused on incremental generation events that increased the text at the end. After each keystroke, if the number of final characters at the end of the text increased relative to the previous event, we labeled the event as an end-of-text generation (i.e., an incremental commit). Events occurring in the middle of the text or involving deletions were excluded from this analysis. For each end-of-text commit, we computed (a) the interval between the final Pinyin-letter keystroke and the appearance of the Chinese character, as an indicator of cognitive effort in candidate recognition and decision making, and (b) the interval between consecutive commits, to capture ideation or planning processes following completion of a production unit.
A data-driven pause threshold of 1,411 ms (μ + 1σ of the pooled inter-event interval distribution, excluding task-unrelated interruptions >10 s) was applied consistently across both tasks.
A key challenge in Pinyin analysis is how to segment Pinyin input in a way that captures individual input habits. To address this, we developed a Pinyin–character alignment model that maps the Pinyin input stream to the final committed Chinese characters(Figure 2). The alignment uses character-commit events as boundaries. For each commit, the model traces backward to retrieve the most recent contiguous sequence of Pinyin letters and treats it as a candidate Pinyin string, which is then enumeratively matched to the committed character(s).

Based on the selected alignment, we identify and count the use of different Pinyin input strategies (e.g., full Pinyin, initial-letter abbreviation, and prefix-based abbreviations, see Table 1). To control for differences in task duration and text length, strategy usage, pause frequency, and related indicators are normalized by the total number of committed characters, yielding proportion-based measures.
Strategy Type | Target Characters | Input Letters | Description |
|---|---|---|---|
Full Pinyin | 知识 | zhishi | All letters of the full Pinyin spelling for each syllable are entered. |
Prefix-based abbreviation | 知识 | zhsh | A reduced prefix form is entered, keeping multiple leading letters from each syllable. |
Initial-letter abbreviation | 知识 | zs | Only the first letter of each Pinyin syllable is entered. |
The alignment model successfully resolved 87.75% of all commit events (alignment coverage rate), yielding a mean of 2.35 committed characters per commit event. The remaining 12.25% of events, in which no valid Pinyin string could be retrieved within the backward-search window or in which the constraint-based filtering could not uniquely resolve the alignment, were excluded from strategy-level analyses but retained for timing-based analyses (e.g., inter-key intervals and pause rates), which do not require Pinyin segmentation.
Among successfully aligned commits, the strategy classification yielded the following aggregate distribution: full Pinyin (M = 77.69%), initial-letter abbreviation (M = 12.86%), prefix-based abbreviation (M = 7.83%), and grouped input with backspace correction (M = 1.62%). The last category comprised events in which the Pinyin buffer was interrupted by mid-buffer deletions prior to character commitment; because such sequences could not be unambiguously attributed to a single input strategy, they were excluded from subsequent strategy-level comparisons. No events fell into the residual "other" category (M = 0.00%), indicating that the constraint-based filtering algorithm exhaustively accounted for all strategy types within the aligned portion of the data.
Based on their location in the input stream, we categorized pauses into three types. (1) Within-character pauses are long intervals occurring during Pinyin letter entry and may reflect difficulty in phonological encoding or momentary shifts of attention. (2) Selection pauses are long intervals between completing Pinyin entry and selecting a candidate character, capturing cognitive load associated with orthographic recognition and decision making. (3) Post-commit pauses are long intervals after a character is committed and before the next input event, which may signal planning at the lexical, syntactic, or discourse level. For all three pause types, occurrence rates were normalized by the total number of committed characters (i.e., computed per character).
Beyond backspace rate (immediate deletions), we also computed a broader edit rate, defined as the proportion of all input operations that involved revision actions in the middle of the text (e.g., deletions and substitutions). This measure was used to capture monitoring and revision activity during writing.
For time-based measures (e.g., inter-key interval, selection latency, and post-commit gap), we report both the mean and the median to better represent their typically skewed distributions. For inter-key interval, we additionally computed the standard deviation and the coefficient of variation () to quantify the stability and relative variability of input rhythm.
Comparative Analysis
After extracting Pinyin input process features, we compared the two writing tasks on each metric. All analyses were conducted within participants to control for stable individual differences.
For each metric, we first tested the normality of the paired difference scores using the Shapiro-Wilk test. If the normality test was not significant (p > .05), we used a paired-samples t test to compare task conditions, otherwise Wilcoxon signed-rank test was applied. We controlled the inflation of Type I error due to multiple comparisons using the Benjamini-Hochberg (BH) procedure. Statistical significance was determined based on BH-adjusted p values.
Correlating Analysis
The writing quality of Task 2 was assessed using Nussbaum’s holistic argumentation rubric, with scores ranging from 0 to 7. All essays were scored independently by two trained raters. Inter-rater reliability was evaluated using Kendall’s tau (τ) and linearly weighted Cohen’s kappa (κ). Agreement was acceptable to high (τ = 0.840, κ = 0.671), indicating that the scoring procedure was reliable. The holistic argumentation scores derived from Task 2 yielded a mean of 4.51 (SD = 1.31, range = 1.50–6.50). For subsequent analyses, we used the mean of the two raters’ scores.
To examine the predictive value of Pinyin input features for writing quality, we related 14 key process metrics extracted from Task 2 to the final argumentation scores. A Shapiro-Wilk test indicated that the score distribution did not significantly deviate from normality (W = 0.953, p = 0.151), supporting the use of Pearson correlation analysis.
RESULTS
Comparative Analysis of Pinyin Input Behaviors Across Tasks
The study first compared the characteristics of participants’ Pinyin input behaviors between the transcription task (Task 1) and the argumentative writing task (Task 2). As shown in Table 2, the shift in task type exerted a significant impact on participants’ input rhythm, cognitive processing latency, and input strategies, providing preliminary support for Hypothesis 1 (H1).
Metric | Test Type | padj | d | r | z |
|---|---|---|---|---|---|
Full Input Ratio | T-test | 0.214 | -0.233 | ||
Initial Input Ratio | T-test | 0.077 | -0.334 | ||
Prefix Input Ratio | Wilcoxon | 0.030 | 0.459 | -2.676 | |
Mean IKI | Wilcoxon | <0.001 | 0.872 | -5.086 | |
SD of IKI | Wilcoxon | <0.001 | 0.872 | -5.086 | |
CV of IKI | T-test | <0.001 | 1.577 | ||
Backspace Rate | Wilcoxon | 0.027 | 0.682 | -3.975 | |
Mean Selection Latency | T-test | <0.001 | -1.252 | ||
Median Selection Latency | T-test | <0.001 | -1.477 | ||
Median Post-commitment Gap | Wilcoxon | <0.001 | 0.849 | -4.949 | |
Global Edit Rate | Wilcoxon | 0.030 | 0.394 | -2.299 | |
Within-character Pause Rate | Wilcoxon | 0.454 | 0.216 | -1.257 | |
Selection Pause Rate | T-test | <0.001 | -0.798 | ||
Post-commit Pause Rate | T-test | 0.370 | 0.165 |
|
|
At the physical keystroke level, a systemic decline in input fluency was observed when participants transitioned from simple character transcription to complex argumentative writing. Statistical results indicated that the mean Inter-Key Interval (Mean IKI) and its standard deviation (SD of IKI) were significantly higher in Task 2 than in Task 1 (Z = -5.086, padj < .001, r = .872). More importantly, the Coefficient of Variation of IKI (CV of IKI), which measures the relative stability of input rhythm, increased drastically in the writing task (padj < .001, Cohen's d = 1.577).
Changes in cognitive-layer metrics further revealed the cognitive load imposed by the writing task. During the candidate character recognition phase, both the mean and median Selection Latency were significantly longer in the writing task compared to the transcription task (padj < .001). Simultaneously, the Selection Pause Rate increased significantly (padj < .001). Furthermore, the Median Post-commit Gap was significantly extended in Task 2 (Z = -4.949, padj < .001, r = .849).
Regarding input strategies and monitoring, participants exhibited distinct compensatory and regulatory tendencies. While the ratios of full-pinyin and initial-pinyin strategies showed no significant changes, the Prefix Input Strategy Ratio rose significantly in the writing task (Z = -2.676, padj = .030, r = .459). Meanwhile, both the Backspace Rate (padj = .027) and the Global Edit Rate (padj = .030) were significantly higher in Task 2 than in Task 1. This indicates that the writing task triggered more frequent evaluation and revision activities\.
Correlation Between Pinyin Input Features and Argumentative Writing Quality
To address RQ3, we conducted Pearson correlation analyses between the 14 key process features extracted from the creative writing task (Task 2) and the holistic argumentation scores. The results revealed that cognitive-layer pause metrics were the most robust predictors of writing quality, while physical-layer fluency and input strategies showed no significant linear relationship with the final product's quality (see Table 3).
Metrics | r | p | 95%CL |
|---|---|---|---|
Full Input Ratio | -0.127 | 0.480 | [-0.451, 0.226] |
Initial Input Ratio | 0.249 | 0.162 | [-0.103, 0.546] |
Prefix Input Ratio | -0.204 | 0.255 | [-0.511, 0.150] |
Mean IKI | -0.123 | 0.496 | [-0.447, 0.230] |
SD of IKI | -0.058 | 0.750 | [-0.393, 0.291] |
CV of IKI | 0.016 | 0.928 | [-0.329, 0.358] |
Backspace Rate | 0.036 | 0.844 | [-0.311, 0.374] |
Mean Selection Latency | -0.036 | 0.842 | [-0.375, 0.311] |
Median Selection Latency | -0.114 | 0.528 | [-0.440, 0.239] |
Mean Post-commit Gap | 0.018 | 0.922 | [-0.328, 0.359] |
Global Edit Rate | 0.175 | 0.330 | [-0.179, 0.489] |
Within-character Pause Rate | -0.482** | 0.005 | [-0.708, -0.166] |
Selection Pause Rate | -0.030 | 0.870 | [-0.369, 0.317] |
Post-commit Pause Rate | -0.465** | 0.006 | [-0.697, -0.145] |
Statistical analysis identified two critical cognitive-layer indicators that were significantly and negatively correlated with argumentation scores. Specifically, the Within-character Pause Rate (r = -.482, p = .005, 95% CI [−.708, −.166]) and the Post-commit Pause Rate (r = -.465, p = .006, 95% CI [−.697, −.145]) showed moderate negative correlations with writing quality.
Notably, physical-layer metrics such as Mean IKI (r = -.123, p = .496, 95% CI [-0.447, 0.230]) and Median Selection Latency (r = -.114, p = .528, 95% CI [-0.440, 0.239]) failed to reach statistical significance. This decoupling indicates that the mere speed of finger movement or character selection does not inherently translate to superior intellectual output in Chinese writing. Furthermore, input strategy ratios, such as the Initial Input Ratio (r = .249, p = .162, 95% CI [-0.103, 0.546]) and Prefix Input Ratio (r = -.204, p = .255, 95% CI [-0.511, 0.150]), were also not significantly associated with scores.
Table 4 provides a compact summary of hypothesis testing results. H1a and H1b were fully supported, showing greater timing variability and longer cognitive processing intervals in writing than copying. H1c received partial support. Revision-related behaviors (backspace rate and global edit rate) were significantly higher in the writing task, while only one of the strategic input indicators (prefix-based input ratio) showed a significant increase. Other strategy ratios did not differ reliably across tasks. H2 was supported: only cognitively grounded pause indicators were significantly associated with writing quality, whereas low-level motor timing metrics were not.
Indicators | Key evidence | Verdict | |
|---|---|---|---|
H1a | CV of IKI | padj < .001, d = 1.577 | Supported |
H1b | Selection latency | padj < .001, d = -1.252 | |
H1b | Post-commit gap | padj < .001, r = 0.849 | Supported |
H1c | Backspace rate | padj = .027, r = 0.682 | Supported |
H1c | Global edit rate | padj = .030, r = 0.394 | Supported |
H1c | Full Input Ratio | padj = .214, d = -0.233 | Not Supported |
H1c | Initial Input Ratio | padj = .077, d = -0.334 | Not Supported |
H1c | Prefix Input Ratio | padj = .030, r = 0.459 | Supported |
H2 | Within-character pause rate vs. Score | r = −.48, p < .01 | Supported |
H2 | Post-commit pause rate vs. Score | r = −.47, p < .01 | Supported |
DISCUSSION
The present study investigated the characteristics of Chinese Pinyin input behaviors and their relationship with writing quality through the lens of a Dual-layer Rhythm Model. Utilizing fine grained keystroke logging and feature engineering techniques, we captured the nuanced shifts in input patterns as participants transitioned from transcription to argumentative writing. The results confirmed that task induced cognitive load significantly alters physical rhythm, cognitive latency, and input strategies. Most crucially, the correlation analysis revealed that cognitive layer pauses metrics, specifically the Within-character Pause Rate and the Post commit Pause Rate, serve as robust predictors of argumentation quality, whereas physical fluency metrics such as Mean Inter Key Interval do not.
A central finding of this study is the significant negative correlation between the Within-character Pause Rate, the Post-commit Pause Rate, and holistic writing scores. From the perspective of feature engineering, these two metrics constitute diagnostic markers for identifying cognitive bottlenecks during the writing process.
The Within-character Pause Rate reflects difficulties during the phonological encoding phase. In Pinyin-mediated writing, the writer must transform abstract mental concepts into specific phonetic sequences. High quality writers exhibit a high degree of automaticity in this phase. Conversely, frequent prolonged pauses during Pinyin spelling among lower scoring writers suggest that lower-level phonology to orthography conversion consumes excessive cognitive resources. Crucially, this relationship may be bidirectional: the intensive cognitive load inherent in high-level argumentative planning can interfere with the execution of motor sequences, causing even potentially automated Pinyin input to stutter. According to Working Memory Theory, such non automated lower-level processing creates cognitive competition, thereby depleting the mental resources available for high level argumentative planning and reasoning.
The Post-commit Pause Rate marks the transition from text generation to higher-level macro-planning. Higher-quality writers typically demonstrate stronger pre-planning capacity or more fluent syntactic production, resulting in fewer disruptions between linguistic units. This finding suggests that typing speed alone is not a reliable indicator of writing quality. Consistent with writing-process and keystroke-logging research, cognitively grounded timing features—such as pause distribution—provide more informative signals of composing quality than raw physical execution speed, especially in Pinyin-mediated Chinese writing.
The observed decoupling between physical layer metrics such as Mean Inter Key Interval and writing scores provides significant insights for educational data mining. This phenomenon suggests that in a population of graduate students, physical execution of Pinyin input has reached a stage of high automaticity, essentially becoming background noise in the writing process.
The true variance in writing expertise is concealed within the decision latencies and pause patterns that occur after the physical keystroke sequences. This validates the necessity of our proposed DLRM. Only by isolating the interference of physical execution can researchers accurately capture the cognitive features truly associated with the quality of thought and argumentation.
From the standpoint of Educational Data Mining, the identification of these two critical pause features provides an empirical foundation for developing process oriented intelligent writing systems. Future automated writing evaluation tools could move beyond product-based scoring to provide real time diagnostics. By monitoring the Within-character Pause Rate and the Post commit Pause Rate, a system could distinguish whether a writer is struggling with spelling retrieval or conceptual planning and offer personalized scaffolding accordingly.
Despite these insights, several limitations should be noted. First, the sample size of 34 and the focus on graduate students limit the generalizability of the findings to younger or less proficient populations. Future research should include diverse groups, such as elementary students or second language learners, to test the model’s robustness.
Second, this study primarily focused on the incremental generation phase. Macro level editing behaviors, such as structural reorganization or cross paragraph revisions, were not explored in depth. Future studies could integrate eye tracking technology with sequential pattern mining to investigate the recursive nature of the writing process, thereby constructing a more comprehensive feature library for Chinese digital writing.
CONCLUSION
This study proposed and validated a Dual-layer Rhythm Model to analyze the cognitive mechanisms underlying Chinese Pinyin input during writing tasks. By leveraging fine grained keystroke logging and data mining techniques, we successfully decoupled the physical execution layer from the cognitive processing layer in digital writing. The results demonstrate that while high level writing tasks induce a systemic decline in physical typing fluency, it is the cognitive layer pauses, specifically those occurring within Pinyin character construction and after character commitment, that serve as significant indicators of writing quality.
The theoretical contribution of this research lies in identifying the unique phonological to orthographical translation costs inherent in Pinyin input, which have been largely overlooked in traditional alphabetic keystroke research. Practically, the identified diagnostic features provide a foundation for developing next generation intelligent writing environments. These systems can transition from static text evaluation to dynamic process monitoring, offering timely interventions when cognitive bottlenecks are detected. Future work will focus on integrating these process features into machine learning models to enhance the predictive accuracy of automated writing assessment systems across diverse learner populations.
Appendix
Appendix A. Transcription Text
假如说逻辑一般被认为是思维的科学,那么,人们对于它的了解是这样的,即:好像这种思维只构成知识的单纯形式;好像逻辑抽去了一切内容,而属于知识的所谓第二组成部分,即质料,必定另有来源;好像完全不为这种质料所依赖的逻辑,因而只能提供真正知识的形式条件,而不能包含实在的真实本身,也不能是达到实在的真理的途径,因为真理的本质的东西,内容,恰恰在逻辑以外。
Appendix B. Argument Prompt
The following is a memorandum from the business manager of a television station. “Over the past year, our late-night news program has devoted increased time to national news and less time to weather and local news. During this period, most of the complaints received from viewers were concerned with our station’s coverage of weather and local news. In addition, local businesses that used to advertise during our late-night news program have canceled their advertising contracts with us. Therefore, in order to attract more viewers to our news programs and to avoid losing any further advertising revenues, we should expand our coverage of weather and local news on all our news programs.”
Write a response in which you examine the stated and/or unstated assumptions of the argument. Be sure to explain how the argument depends on these assumptions and what the implications are for the argument if the assumptions prove unwarranted.
Ethics and Consent to Participate
Formal ethical approval was not required for this study in accordance with national and institutional guidelines, as determined by the Institutional Review Board of the University of Science and Technology of China. The study involved anonymous data collection and posed no foreseeable risk to adult participants. Participation was voluntary, and informed consent was obtained from all participants prior to data collection. If required, documentation confirming the exemption from IRB review can be provided.
ACKNOWLEDGMENTS
This research was also supported by the advanced computing resources provided by the Supercomputing Center of the USTC.
REFERENCES
- Anderson, J.R. 1982. Acquisition of cognitive skill. Psychological review. 89, 4 (1982), 369. https://doi.org/10.1037/0033-295X.89.4.369.
- Anderson, J.R. 2014. Rules of the mind. Psychology Press.
- Baddeley, A. 2003. Working memory: looking back and looking forward. Nature reviews neuroscience. 4, 10 (2003), 829–839. https://doi.org/10.1038/nrn1201.
- Baddeley, A.D. 2017. Modularity, working memory and language acquisition. Second language research. 33, 3 (2017), 299–311. https://doi.org/10.1177/0267658317709852.
- Baddeley, A.D. 2015. Working memory in second language learning. Working memory in second language acquisition and processing. (2015), 17–28.
- Chen, J., Luo, R. and Liu, H. 2017. The Effect of Pinyin Input Experience on the Link Between Semantic and Phonology of Chinese Character in Digital Writing. Journal of Psycholinguistic Research. 46, 4 (Aug. 2017), 923–934. https://doi.org/10.1007/s10936-016-9470-y.
- Chenoweth, N.A. and Hayes, J.R. 2001. Fluency in Writing: Generating Text in L1 and L2. Written Communication. 18, 1 (Jan. 2001), 80–98. https://doi.org/10.1177/0741088301018001004.
- Conijn, R., Cook, C., van Zaanen, M. and Van Waes, L. 2022. Early prediction of writing quality using keystroke logging. International Journal of Artificial Intelligence in Education. 32, 4 (2022), 835–866. https://doi.org/10.1007/s40593-021-00268-w.
- Conijn, R., Roeser, J. and Van Zaanen, M. 2019. Understanding the keystroke log: The effect of writing task on keystroke features. Reading and Writing. 32, 9 (2019), 2353–2374. https://doi.org/10.1007/s11145-019-09953-8.
- Correnti, R., Matsumura, L.C., Wang, E.L., Litman, D. and Zhang, H. 2022. Building a validity argument for an automated writing evaluation system (eRevise) as a formative assessment. Computers and Education Open. 3, (2022), 100084. https://doi.org/10.1016/j.caeo.2022.100084.
- Crotti, R., Denaro, G., Du, Z. and MartÃn, R.M. 2026. Hylog: A Hybrid Approach to Logging Text Production in Non-alphabetic Scripts. arXiv preprint arXiv:2601.17753. (2026).
- Diamond, A. 2013. Executive functions. Annual review of psychology. 64, 1 (2013), 135–168. https://doi.org/10.1146/annurev-psych-113011-143750.
- Ericsson, K.A. and Simon, H.A. 1980. Verbal reports as data. Psychological review. 87, 3 (1980), 215. https://doi.org/10.1037/0033-295X.87.3.215.
- Guan, C.Q., Liu, Y., Chan, D.H.L., Ye, F. and Perfetti, C.A. 2011. Writing strengthens orthography and alphabetic-coding strengthens phonology in learning to read Chinese. Journal of Educational Psychology. 103, 3 (Aug. 2011), 509–522. https://doi.org/10.1037/a0023730.
- Ho, C.S.-H. and Bryant, P. 1997. Learning to read Chinese beyond the logographic phase. Reading research quarterly. 32, 3 (1997), 276–289. https://doi.org/10.1598/RRQ.32.3.3.
- Huang, C.-R., Lin, Y.-H., Chen, I.-H. and Hsu, Y.-Y. 2022. The Cambridge handbook of Chinese linguistics. Cambridge University Press.
- Kandel, S. and Perret, C. 2015. How do movements to produce letters become automatic during writing acquisition? Investigating the development of motor anticipation. International Journal of Behavioral Development. 39, 2 (2015), 113–120. https://doi.org/10.1177/0165025414557532.
- Kellogg, R.T. 1996. A MODEL OF WORKING MEMORY IN WRITING. The science of writing: Theories, methods, individual differences, and applications. Lawrence Erlbaum Associates, Inc. 57–71.
- Kellogg, R.T. 2008. Training writing skills: A cognitive developmental perspective. Journal of Writing Research. 1, 1 (June 2008), 1–26. https://doi.org/10.17239/jowr-2008.01.01.1.
- Leijten, M. and Van Waes, L. 2013. Keystroke Logging in Writing Research: Using Inputlog to Analyze and Visualize Writing Processes. Written Communication. 30, 3 (July 2013), 358–392. https://doi.org/10.1177/0741088313491692.
- Li, T., Yu, L., Zhou, L. and Wang, P. 2023. Using less keystrokes to achieve high top-1 accuracy in Chinese clinical text entry. Digital Health. 9, (2023), 20552076231179027. https://doi.org/10.1177/20552076231179027.
- Logan, G.D. 1988. Toward an instance theory of automatization. Psychological review. 95, 4 (1988), 492. https://doi.org/10.1037/0033-295X.95.4.492.
- Lu, X. and Révész, A. 2021. Revising in a non-alphabetic language: The multi-dimensional and dynamic nature of online revisions in Chinese as a second language. System. 100, (Aug. 2021), 102544. https://doi.org/10.1016/j.system.2021.102544.
- Lyu, B., Lai, C., Lin, C.-H. and Gong, Y. 2021. Comparison studies of typing and handwriting in Chinese language learning: A synthetic review. International Journal of Educational Research. 106, (2021), 101740. https://doi.org/10.1016/j.ijer.2021.101740.
- McCutchen, D. 2000. Knowledge, Processing, and Working Memory: Implications for a Theory of Writing. Educational Psychologist. 35, 1 (Jan. 2000), 13–23. https://doi.org/10.1207/S15326985EP3501_3.
- Oxley, J. and Ma, Y. 2020. Considerations for Chinese text input methods in the design of speech generating devices: a tutorial. Clinical Linguistics & Phonetics. 34, 4 (2020), 366–387. https://doi.org/10.1080/02699206.2019.1652934.
- Perfetti, C.A., Liu, Y. and Tan, L.H. 2005. The lexical constituency model: some implications of research on Chinese for general theories of reading. Psychological review. 112, 1 (2005), 43. https://doi.org/10.1037/0033-295X.112.1.43.
- Perfetti, C.A. and Tan, L.H. 1998. The time course of graphic, phonological, and semantic activation in Chinese character identification. Journal of Experimental Psychology: Learning, Memory, and Cognition. 24, 1 (1998), 101. https://doi.org/10.1037/0278-7393.24.1.101.
- Shen, H.H. and Bear, D.R. 2000. Development of orthographic skills in Chinese children. Reading and Writing. 13, 3 (2000), 197–236. https://doi.org/10.1023/A:1026484207650.
- Sun, X. and Huo, X. 2022. Word-Level and Pinyin-Level Based Chinese Short Text Classification. IEEE Access. 10, (2022), 125552–125563. https://doi.org/10.1109/ACCESS.2022.3225659.
- Sweller, J. 2011. Cognitive load theory. Psychology of learning and motivation. Elsevier. 37–76.
- Talebinamvar, M. 2022. Clustering students’ writing behaviors using keystroke logging: a learning analytic approach in EFL writing. Language Testing in Asia. 12, 1 (2022), 6.
- Van Waes, L. and Leijten, M. 2015. Fluency in Writing: A Multidimensional Perspective on Writing Fluency Applied to L1 and L2. Computers and Composition. 38, (Dec. 2015), 79–95. https://doi.org/10.1016/j.compcom.2015.09.012.
- Wang, J., Zhai, S. and Su, H. 2001. Chinese input with keyboard and eye-tracking: an anatomical study. Proceedings of the SIGCHI conference on Human factors in computing systems (2001), 349–356.
- Wang, M., Koda, K. and Perfetti, C.A. 2003. Alphabetic and nonalphabetic L1 effects in English word identification: A comparison of Korean and Chinese English L2 learners. Cognition. 87, 2 (2003), 129–149.
- Wang, X. and Wang, J. 2024. Comparing Chinese L2 writing performance in paper-based and computer-based modes: Perspectives from the writing product and process. Assessing Writing. 61, (2024), 100849. https://doi.org/10.1016/j.asw.2024.100849.
- Zhai, S. and Kristensson, P.O. 2012. The word-gesture keyboard: reimagining keyboard interaction. Communications of the ACM. 55, 9 (2012), 91–101. https://doi.org/10.1145/2330667.2330689.
- Zhou, X. and Marslen-Wilson, W. 1999. Phonology, orthography, and semantic activation in reading Chinese. Journal of Memory and Language. 41, 4 (1999), 579–606. https://doi.org/10.1006/jmla.1999.2663.
© 2026 Copyright is held by the author(s). This work is distributed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0) license.