2025-10-10
This chapter introduces a study which investigates the features of connector use in Chinese nonEnglish major college students’ writing based on a home-developed learner corpus (3,908,816 tokens). The research findings reveal that Chinese college students use more one-word connectors than multiword ones to express the meaning of enumeration and addition, and that connectors are usually placed in simple declarative sentence order with inverted sentence order or complex sentence patterns rarely used. With reference to the English Grammar Profile (EGP), the criterial features of grammatical use based on the CEFR, it can be found that students’ use of connectors spreads across levels, with most connectors clustered on lower levels. It is expected that the current empirical study can inform scale descriptors and criterial features of the Cohesion Competence of the China’s Standards of English Language Ability (CSE).
Drawing on a previous large-scale study examining the reactions of past candidates to the use of online invigilation – or online ‘proctoring’ (OLP) – in the delivery of high-stakes English language examinations (Coniam et al., 2021), this paper reports the responses of the subset of China candidates in the sample. The paper first sets the scene in terms of the gradual and then accelerated move from face to face to online modes of delivery. It explores the challenges and benefits that both modes offer, in terms of accessibility, fairness, security and cheating. Detail is then presented from the survey exploring the reactions to and perceptions of OLP by the China respondents (N=64), comparing this sample with the larger world-wide sample, all of whom had taken an English language examination via OLP. A strong endorsement by the China cohort of OLP was generally recorded. Feedback revealed that respondents perceived OLP to be a more personal as well as a more efficient way of taking a test. The results are indicative of a broad acceptance of OLP, pointing to strong future uptake of the OLP mode of test delivery.
This article describes the importance of including the construct of interactional competence in speaking assessments, drawing mainly from the literature in the field of language testing. The coconstruction of meaning and the shared nature of the interaction are seen to be operationalised in an optimal manner using the role play task. The effect of the task is explored through the perspective of the LanguageCert International ESOL Speaking exams, which are used as examples to demonstrate the issues of scalability, discriminability, score separability issues, and the so-called interlocutor effect. Further research and technological innovations will assist in defining and scrutinising the aspects of interactional competence that can be reliably measured.
The last two decades have witnessed language standards frameworks serving as guidance and the metric for the design and enactment of large-scale language assessment. In the United States (U.S.), federal legislation has also been a driving force in dictating the parameters for such an alliance. This article provides an illustrative example of how the thinking of a consortium of U.S. states, territories, and federal agencies has evolved in its conceptualization of language standards frameworks, working in tandem with the development of language assessment to ensure robust alignment between the two. Theoretical and historical foundations underpin the discussion while the components of the language frameworks with their close relation to language assessment are presented to substantiate validity claims between standards frameworks and assessment.
Commonly, models of reading emphasize the relevance of cognitive processes and metacognitive strategies when it comes to reading comprehension. Although these models are derived from research into first language (L1) reading, they often serve as the basis for operationalizing foreign language (FL) reading comprehension tests. This is also true for the FL Reading Test (E8 Reading Test) administered in 2013 and 2019 in Austria. However, little is known about test takers’ behaviour regarding their cognitive processes and metacognitive strategies when taking the test. Undoubtedly, such information could be useful for the FL classroom and to inform foreign language reading instruction as well as test development. Therefore, this paper investigates which cognitive processes and metacognitive strategies students (n=106) apply when taking the E8 Reading Test. Data was collected using a reflective questionnaire. The results show that compared to weaker readers, strong readers apply the expected cognitive processes more frequently. There is no statistically significant difference, however, between stronger and weaker readers regarding metacognitive strategies applied. Furthermore, the data revealed that some good readers arrive at the correct answer by making use of different/other cognitive processes and metacognitive strategies than expected. These findings emphasize the importance of explicit, guided foreign language reading instruction focussing not only on the product of reading (comprehension) but also on the processes involved. However, more research is needed to better understand what the absence or presence of skills and strategies mean regarding individual learner abilities.
Recently, there has been a marked increase in language testing research involving eye-tracking. It appears to offer a useful methodology for examining cognitive validity in language tests, i.e., the extent to which the mental processes that a language test elicits from test takers resemble those that they would employ in the target language use domains. This article reports on a recent study which examined reading processes of test takers at different proficiency levels on a reading proficiency test. Using a mixed-methods approach, the study collected cognitive validity evidence through eyetracking and stimulated recall interviews. The study investigated whether there are differences in reading behaviour among test takers at CEFR B1, B2 and C1 levels on an online reading task. The main findings are reported and the implications of the findings are discussed to reflect on some fundamental questions regarding the use of eye-tracking in language testing research.
Assessment reform has been one of the major concerns of English teaching reform in senior secondary vocational schools in China. The Achievement Test on English Teaching in senior secondary vocational schools in Beijing was launched in June 2015 to better assess the effects of English teaching and learning in senior secondary vocational schools. Students who have completed the basic module of the English course take the test. This paper reviews briefly the history of achievement test, discusses the research on washback at home and abroad, and analyses the structure and components of the achievement test carried out in Beijing. It focuses on the washback effect of the achievement test in particular by comparing and contrasting the expected effect of the test with the real feelings and behaviors of teachers, thus providing implications for English teaching and learning. This paper concludes with the observation that while the achievement test has a number of positive effects, efforts nonetheless need to be made to prevent any potential negative effects.
This paper addresses the topic of English assessment in senior secondary vocational schools in the Chinese mainland. Studies are reviewed concerning the implementation of the concept of assessment proposed in the national syllabus issued in 2009. With reference to the concept of assessment in the 2009 Syllabus and the 2020 Curriculum, an online questionnaire survey was conducted to investigate the practice of assessment in senior secondary vocational schools in China. Sample tests used in senior secondary vocational schools were drawn upon to provide evidence for the practice of assessment. Findings from this study indicate that a shift has been made from traditional assessment of learning to assessment for and as learning. The integrated model of assessment proposed in the Syllabus is commonly practised. Some other models like multiple assessment and credit substitution assessment are also adopted in some middle schools. Diverse evaluation criteria are adopted, multiple methods of assessment are employed, students and stakeholders are participating in assessment as agents. Teachers are using formative assessment to support teaching and learning. Innovation can also be found in summative assessment, in which performance tasks are designed, authentic materials adopted, and real-life/vocational contexts provided for measuring students’ ability to use language to do things.
This research presents the development of an online speaking test of English for students at the end of primary and beginning of secondary school education in state schools in Uruguay. Following the success of the Plan Ceibal one computer-tablet per child initiative, there was a drive to further utilize technology to improve the language ability of students, particularly in speaking, where the majority of students are at CEFR levels pre-A1 and A1. The national concern over a lack of spoken communicative skills amongst students led to a decision to develop a new speaking test, specifically tailored to local needs. This paper provides an overview of the speaking test development and validation project designed with the following objectives in mind: to establish, track, and report annually learners’ achievements against the Common European Framework of Reference for Languages (CEFR) targeting CEFR levels pre-A1 to A2, to inform teaching and learning, and to promote speaking practice in classrooms. Results of a three-phase mixed-methods study involving small-scale and large-scale trials with learners and examiners as well as a CEFR-linking exercise with expert panelists will be reported. Different sources of evidence will be brought together to build a validity argument for the test. The paper will also focus on some of the challenges involved in assessing young learners and discuss how design decisions, local knowledge and expertise, and technological innovations can be used to address such challenges with implications for other similar test development projects.
This paper reports on an exploratory comparability study between the Common European Framework of Reference for Languages (CEFR) and the China Standards of English (CSE). Established equivalences are exhibited via the LanguageCert Test of English of reading and language use for the CEFR and a comparable test of reading and language use produced by a top-tier China university. In the study, a large sample of test takers took part, first sitting the two comparable tests of reading and language use, and subsequently completing a number of self-assessment Can-Do statements related to the CEFR and the CSE. Validity of the dataset was established by linking both tests and sets of self-assessments to a single frame of reference using a third test whose robustness and values had been previously established. While there were some divergences between how the two frameworks aligned – more notably towards the lower ends of the scales – correspondences which emerged between the CEFR and CSE frameworks were broadly in accordance with those reported in other studies referenced in the current paper. The current study therefore sets the groundwork for determining the correspondence between LanguageCert Tests, aligned to the CEFR, and the CSE.
This paper reports on the use of externally-referenced anchoring by LanguageCert as a methodology for calibrating language test materials and aligning test forms. The datasets used are taken from tests at each of the six levels of LanguageCert IESOL suite, all of which have been aligned to the CEFR through expert judgement. We illustrate in this paper the extent to which externally-referenced anchoring, using Item Response Theory (IRT) but based on expert judgement, can be used as an effective, reliable and valid methodology. The approach is based on the premise that successful anchoring may be achieved by reference to well-targeted, expertly-written test forms aligned to the underlying traits of a particular CEFR level by expert judgement and verified through the use of IRT. This study focuses on the analysis of 18 LanguageCert test forms, three at each CEFR level. The LanguageCert Item Difficulty (LID) scale, which underlies all LanguageCert test materials, is linked empirically to the CEFR, and each test was placed on the LID scale based at the midpoint of its distribution. This midpoint setting was then set as the externally-referenced anchor for a given CEFR level. The findings of this study indicate that, while the match between the distribution of items in the selected LanguageCert IESOL tests and the LID scale was not perfect, in general, a relatively close match between the items in the tests and the LID scale was found and, as a consequence, the corresponding CEFR level. For each test, most of the items fell between the 25th and 75th percentile of any given level: this range representing the lower and upper bounds of LID scale values for each CEFR level. These results demonstrate that LanguageCert IESOL test items are well set and appropriately positioned at respective CEFR levels on the basis of expert judgement. The study illustrates that externally-referenced anchoring based on expert judgement may be used as a methodology for aligning test forms to an external frame of reference, in this case the CEFR.
This paper reports on a study of the training and standardisation of examiners who mark LanguageCert’s International ESOL (IESOL) suite of English language tests linked to the Common European Framework of Reference (CEFR). Subjects in the study were a set of examiners (N=27) who had been marking LanguageCert’s IESOL Writing tests across the six CEFR levels. The focus of the study was on the consistency of marking in terms of severity within and across the six tests that the examiners mark.Correlations between examiner person measures across all six tests indicated that examiners were broadly consistent across tests, with examiner person measures generally correlating highly with their "partner" test: A1 with A2, C1 with C2, and B1 with B2 tests. LanguageCert examiners – who undergo careful training and standardisation – may therefore be seen to mark consistently and accurately across a range of ability levels.
In China, vocational and technical colleges account for more than half of its institutions of higher education. For most students in these colleges, English is compulsory. This creates one of China’s largest English learning bodies. The paper deals with the Practical English Test for Colleges (PRETCO), a nation-wide standardized English test specially designed for non-English majors of vocational and technical colleges. Specifically, it first focuses on its purpose, length, and administration. Then, it describes its development and introduces its components and format at length. It also describes test takers’ performance in the test. Finally, it provides an overall evaluation of the test.
8-72 characters
1 day ago
Your contribution has been approved successfully
12 hours ago
Your submission has failed to be reviewed; Failure reason: the