1. Introduction and problem
In a secondary writing workshop, the learning environment assembles four objects when it returns a draft. An automated comment bank marks “develop the thesis” and “vary the vocabulary.” In the margin, a language-model rubric score assigns a level to coherence. Below, a similarity dashboard colors overlaps with a corpus. A chatbot invites the student to “write better” the flagged paragraph. The contract calls the bundle formative assessment and leaves the teacher a button to publish the return. The scene is not field data for this article: it is the display ceiling the market already treats as craft. Nothing is written, before use, about which learning goal authorizes that return. There is no interpretive criterion the teacher can revise, no dialogue with whoever wrote the text, no revision cycle, and no name for who decides progress.
The thesis does not fit in the short title. Those objects do not constitute formative assessment if the teacher-mediation protocol is missing. The craft is the judgment that chooses which feedback matters, when to return it, and how to turn it into learning evidence. Davies (2026) situates generative AI for returning formative comments on writing drafts: the object is feedback on the draft, not an equivalence between an automated bank and a formative practice. Pack, Barrett, and Escalante (2024) examine the validity and reliability of LLM automated essay scoring of English-language-learner writing: the object is the score, not the learning cycle. Tate et al. (2024) ask whether AI can provide useful holistic essay scoring: usefulness of the score is not, by itself, mediation. Status: generative-feedback practice in one case; empirical scoring findings in the others. None proves that the automated comment “already assesses formatively.” They ask whether a score was measured or a craft of return was exercised.
The text rests on three distinctions already made in this series and does not remake them. On 16 September 2026 the harness was understood as the layer that turns generative capacity into work: bounded tools, state, verifiers, a stop. On the 18th, agent evaluation measures whether that work is reliable against a criterion external to the model’s prose. On the 22nd, human-in-the-loop supervision governs the copilot: acceptance criterion, traces, veto, attribution, and an escalation threshold. None of those three layers is the formative evaluative judgment of someone who has a group, a subject, and a learning cycle. A system can produce comments and not mediate. It can score with scoring reliability and not form. It can be governed as a copilot and still not return learning evidence. Design inference, not a finding of our own: without teacher mediation there is a feedback product and there is a metric; there is no formative assessment.
Six illegitimate substitutions organize the problem. Treating GenAI feedback as formative assessment: Banihashem et al. (2024) compare sources —peer or AI—; the source is not the protocol. Treating AES as teacher judgment: Li and Liu (2024) score non-native Japanese; Choi et al. (2026) make the anchor the key to AES by prompting; Kim and Holden (2026) contrast API and GUI interfaces. The score is not learning: Pan (2025) indexes generative feedback to a rubric; indexing does not sign the cycle. Treating peer-with-AI as a substitute for the teacher: Guo et al. (2024) study AI-supported peer feedback; Bauer et al. (2023) want NLP to support the exchange, not to replace it. Confusing prompt literacy with evaluative literacy: Kokoç et al. (2026), Yang and Banks (2025), and Zhang et al. (2025) map teacher, teacher-educator, and student AI literacy; knowing how to ask the model is not knowing how to judge progress. Believing that the harness, agent eval, or HITL already mediates: Yan et al. (2024) map practical and ethical challenges; Alfredo et al. (2024) and Capel and Brereton (2023) recall that human-centeredness is not the catalog. This article separates artifacts and mediation and proposes four tests that are not a standard. No Ns, d, r, AUC, percentages, or DOIs are invented.
2. State of the art: situated practice vs artifact
The discourse of “formative assessment with AI” crushes four strata. The first is LLM AES as a problem of validity, reliability, anchor, and scoring interface (Pack et al., 2024; Li and Liu, 2024; Choi et al., 2026; Kim and Holden, 2026; Tate et al., 2024). The second is GenAI as a feedback source on drafts or as rubric-indexed feedback (Davies, 2026; Pan, 2025; Banihashem et al., 2024). The third is peer feedback with AI or NLP support, and the training of whoever comments (Guo et al., 2024; Bauer et al., 2023; Xu et al., 2024; Prilop and Weber, 2023). The fourth is literacy, human-centered assessment, and collaborative competence (Kokoç et al., 2026; Yang and Banks, 2025; Zhang et al., 2025; Alfredo et al., 2024; Yan et al., 2024; Capel and Brereton, 2023; Balducci, 2024; Cui et al., 2026; Goslen et al., 2025). The artifact lives in the grader advertisement. Situated practice lives in the goal, the criterion, the dialogue, and the name of who judges.
In scoring, Pack et al. (2024) saturate the ceiling of validity and reliability when an LLM scores English-language-learner writing: an empirical AES finding, not formative assessment. Li and Liu (2024) extend the application to non-native Japanese: the object remains scoring. Choi et al. (2026) argue that the anchor is the key to accessible AES through prompting: the anchor is a scoring technique, not a dialogued interpretive criterion. Kim and Holden (2026) contrast API and GUI: changing the scoring surface does not install a revision cycle. Tate et al. (2024) ask about the usefulness of holistic scoring: “useful” does not authorize reading the number as learning evidence. Category inference: the AES state of the art authorizes a demand for score validity; it does not authorize selling the score as formative mediation.
In feedback and peer work, Davies (2026) treats GenAI as a formative return on drafts: it names the use; it does not equate an automated bank with classroom craft. Pan (2025) indexes feedback to a rubric in medical English: a declared strategy, not proof that a self-applied rubric replaces the teacher. Banihashem et al. (2024) oppose sources —peer or AI—: comparing provenance does not close the protocol. Guo et al. (2024) trial AI-supported peer feedback on comment quality and writing: a finding of support for the peer, not of substitution of judgment. Bauer et al. (2023) propose a framework: NLP holds up peer feedback; it does not make it formative by generating text. Xu et al. (2024) study visualization of peer-assessment for pre-service teachers: seeing feedback is not mediating it. Prilop and Weber (2023) show that expert feedback, in a digital video-based training, bears on beliefs and comment quality: the expert remains a mediator. Inference: there is situated practice when someone with craft holds the cycle; there is an artifact when the source or the visualization takes that place.
In literacy and human-centered design, Kokoç et al. (2026) review in-service teacher AI literacy: the map is not evaluative literacy. Yang and Banks (2025) locate literacy in elementary teacher educators, not in a prompt workshop. Zhang et al. (2025) validate a student inventory: the construct is not the grading judgment. Alfredo et al. (2024) review human-centred analytics and AI: excluding teachers and students from design feeds mistrust. Capel and Brereton (2023) ask what is human-centered about that label: the map does not authorize calling a grader centered. Balducci (2024) situates assessment in a human-centered education: the act is not exhausted by automation. Cui et al. (2026) assess human–AI collaborative competence in L2 writing: the object is joint competence, not the score. Goslen et al. (2025) generate student plans with an LLM for scaffolding in games: the plan is not formative classroom judgment. Yan et al. (2024) leave open the challenges of grading and feeding back with models. Inference: the state of the art distinguishes scoring, source, peer support, literacy, and framework; the market crushes them.
3. Review method
A narrative critical review was conducted, not a meta-analysis and not a rate of “success of GenAI feedback.” The argument is one of category: what counts as formative assessment when a teacher mediates comments, scores, and chatbots, and what remains an artifact. The window is 2021–2026. The focus is formative assessment with GenAI, AES-LLM, peer+AI or NLP, and teacher and student AI literacy. Where print year diverges from the online year, the print year is cited: Cui et al. (2026) despite a 2025 DOI; Yang and Banks (2025) despite 2024 online; Yan et al. (2024) despite 2023 online. Inclusion: peer-reviewed items with Crossref-verified DOI on 23 September 2026, slot 09:02 America/Mexico_City. Early-childhood axes from this series, product catalogs without a paper, and any figure absent from the sources were excluded. The twenty-one entries in fuentes.md were used. No authors, journals, or DOIs were added.
Each source is marked with one of three statuses. Empirical finding: what was observed in scoring, feedback sources, AI-supported peer work, visualization, or comment training. Framework: a review, agenda, map, or instrument that does not trial the four tests. Design inference: what this article concludes and cannot attribute as a result. Without a study that equates them with formative assessment, the comment bank, the AES-LLM score, the dashboard, the chatbot, the self-applied rubric, and the catalog grader are a category ceiling. There is no PRISMA of our own. The tests in section 7 are hypotheses, not a standard. No Ns, d, r, AUC, interrater percentages, or DOIs are invented.
4. Axis 1. Artifacts do not constitute support
The support this axis denies is not that of an adult in an early-childhood room. It is what a vendor declares when saying the teacher “already has formative assessment” because the system commented on the draft. The GenAI comment bank is the first artifact. Davies (2026) names the use of generative models to return formative feedback on drafts: pasting phrases in the margin can coexist with the absence of a prior goal and of a dialogue. The bank speaks in templates, as if it had already read the curriculum. “Develop the thesis” does not say which thesis, for which criterion, or in which cycle. Restrictive inference: a store of comments is not a formative practice. It can be an input. Displayed as the assessment, it is a ceiling.
The AES-LLM score is not teacher judgment. Pack et al. (2024) fix the problem in the validity and reliability of scoring English-language-learner writing: the finding lives in the score, not in the classroom as a cycle. Li and Liu (2024) apply the same object to non-native Japanese. Choi et al. (2026) shift accessibility toward prompting and the anchor: the anchor calibrates the model; it does not converse with the learner. Kim and Holden (2026) show that the interface —API or GUI— changes the mode of scoring; it does not install attribution. Tate et al. (2024) ask about holistic usefulness: a number that is “useful” for ranking is not evidence that someone learned. Yan et al. (2024) leave automated feedback and grading among the practical and ethical challenges: that a system scores does not mean the teacher has fixed which error is not delegated. Model confidence and interpretive criterion are not convertible.
The similarity dashboard and the “write better” chatbot complete the grammar of display. Xu et al. (2024) study visualization of peer-assessment: seeing data may touch data literacy, motivation, and load; it does not authorize treating a similarity board as mediation. The color of the overlap does not say whether the borrowing is learning, citation, or a shortcut. The chatbot that rewrites the paragraph occupies the place of the next draft: Goslen et al. (2025) generate student plans with an LLM for scaffolding in games —a neighboring object— and recall that producing a plan or a rewrite is not judging progress. Banihashem et al. (2024) can compare AI feedback with peer feedback without making the chatbot the subject of the cycle. The learner who pastes the output into the next submission has not revised: they have substituted. Without dialogue there is no formative assessment; there is cosmetics of the text.
The self-applied rubric and the marketing “AI grader” close the criterion in false. Pan (2025) indexes generative feedback to a rubric in medical English and declares it a formative strategy: indexing is a prompt design; it is not the interpretive criterion revisable by whoever teaches. Balducci (2024) situates assessment in a human-centered education: automating the rubric does not inherit that centering. Capel and Brereton (2023) forbid calling human-centered what only displays a person in the brochure. Alfredo et al. (2024) ask for real control, not a chart that hides who scores. The catalog grader crushes scoring, feedback, and grading into a single word. Category inference: bank, score, dashboard, chatbot, self-applied rubric, and grader share a grammar of substitution. None is authorized by the corpus.
5. Axis 2. The relational craft in the garden
The heading is from the series template. The referent is not. There is no early-childhood classroom. The ground is the relation between a person responsible for a subject and a system that produces comments, scores, rewrites, and rubrics. There is craft when that person sets goals before using the model, holds interpretive criteria they can revise, dialogues with the learner, opens revision cycles, names who decides the grade or progress, and literacizes feedback —their own and the student’s—. Prilop and Weber (2023) show, without speaking of GenAI as a grader, that expert feedback changes the beliefs and comment quality of those in formation: the mediating subject is not downloaded into a template. A comment bank does not recruit that identity.
The first gesture is the learning goal prior to use of the model, not the impression of the demo. The criterion is not born from the output. Balducci (2024) forbids starting from the tool: student assessment, in a human-centered education, is designed as an act, not as a plugin. Davies (2026) can return comments on drafts only if someone already knew what was being learned in that draft. Pan (2025) indexes to a rubric: if it belongs to the subject and precedes the model, it can be an input; if it is born from the prompt, it is an artifact. Zhang et al. (2025) measure student AI literacy, not the goal of the task. Inference: the goal is written as observable performance and as an error that is not delegated. It is not written as a high score.
The second gesture is the interpretive criterion revisable by the teacher. Choi et al. (2026) treat the anchor as the key to AES by prompting: the anchor calibrates the model; the classroom criterion is discussed, adjusted, and taught. Pack et al. (2024) and Li and Liu (2024) can describe the validity of a score without that score being the criterion of this prompt. Kim and Holden (2026) change the scoring interface: the GUI is not the living rubric. Tate et al. (2024) ask about holistic usefulness: an opaque holistic cannot be revised. Alfredo et al. (2024) tie trust to participating in the design. Design inference: there is a criterion when the teacher can say what counts as “sufficient cohesion” in this genre, with this group, and amend that reading after seeing the drafts. If the only criterion is the number the model returned, there is scoring, not mediation.
The third gesture is dialogue with the learner and the revision cycle. Guo et al. (2024) locate AI-supported peer work in feedback quality and in writing: the peer remains an interlocutor. Bauer et al. (2023) want NLP to hold that exchange, not to close it. Banihashem et al. (2024) compare sources: the comparison only matters if someone returns the text to a cycle. Xu et al. (2024) visualize peer-assessment: visualization can inform; it does not speak with the learner. Prilop and Weber (2023) train the comment: training is already a mediated dialogue. Davies (2026) works on drafts: the draft demands a next version, not a verdict. Inference: there is formative assessment when the comment enters a conversation and produces an attributable revision. If the chatbot rewrites and the student submits that version, the cycle has been skipped.
Attribution and feedback literacy go together. Cui et al. (2026) assess human–AI collaborative competence in L2 writing: the framework obliges one to say who does what. That “who” is a neighbor of the attribution of grading judgment, not its substitute. Yan et al. (2024) forbid treating automated grading as an innocuous gesture. Capel and Brereton (2023) forbid hiding the person behind the label. Kokoç et al. (2026) map teacher literacy: knowing how to use the tool is not knowing how to return a judgment. Yang and Banks (2025) place literacy in teacher educators, not in a click. Zhang et al. (2025) give a student instrument: AI literacy is not evaluative literacy. Goslen et al. (2025) recall the neighbor of the generated plan: scaffolding a game does not name who grades. The name of whoever adopts the score, and of whoever decides progress, is written outside the model. The teacher mediates, corrects, or switches off. They are not a customer of a grader.
6. Contrast. Category boundaries
GenAI feedback is not formative assessment. Davies (2026) and Pan (2025) can describe generative returns —on drafts or indexed to a rubric— without authorizing the equation. Banihashem et al. (2024) show that the source can be AI: the source is not the cycle. Calling the automated comment formative and adding “with the teacher in the loop” in a footnote erases the boundary instead of crossing it. Formative assessment denies that the system is the subject of the judgment.
AES is not teacher judgment, and the score is not learning. Pack et al. (2024), Li and Liu (2024), Choi et al. (2026), Kim and Holden (2026), and Tate et al. (2024) construct a scoring problem: validity, reliability, anchor, interface, holistic usefulness. None installs a dialogue or an attribution of the grade. A well-measured AES can be an input to a rubric. Converted into the student’s progress, it has crossed sides. Xu et al. (2024) can visualize peer-assessment without the color being the learning. Restrictive inference: the number speaks of the text against a scoring model; learning speaks of an attributable change, read by someone with a criterion.
Peer-with-AI does not replace the teacher, and prompt literacy is not evaluative literacy. Guo et al. (2024) and Bauer et al. (2023) treat support for the peer as the object: the teacher still designs the task and, if there is a grade, still signs it. Prilop and Weber (2023) need the expert to form the comment. Kokoç et al. (2026), Yang and Banks (2025), and Zhang et al. (2025) disaggregate AI literacy —in-service teacher, educator, student—: “I know how to ask the model” is not knowing what counts as evidence of progress. Finishing a workshop on “I can already generate the rubric” can be functional literacy and remain evaluative illiteracy.
The harness is not formative mediation, agent eval is not a classroom rubric, and HITL is not attribution of the grading judgment. On 16 September the harness produces work with tools, state, verification, and a stop: a bank can be well harnessed and still lack a goal. On the 18th, agent eval measures reliability on a task distribution: being well measured at scoring does not authorize grading where the threshold is not written. On the 22nd, HITL governs the copilot with veto and trace: a comment can be vetoed without anyone having dialogued or named who decides progress. Goslen et al. (2025) illustrate another boundary: generating a student plan is scaffolding, not formative judgment. Cui et al. (2026) assess collaborative competence: they do not equate that competence with the signature of a grade. Yan et al. (2024), Alfredo et al. (2024), Capel and Brereton (2023), and Balducci (2024) forbid closing the problem with a catalog grader. The category is conjunctive: one piece missing and the system falls back into artifact, even if the slide says formative assessment.
7. Four tests of support (not artifact)
What follows is a design inference of this article, anchored in the corpus and not delivered by any source as a standard. It is not a norm, not a validated rubric, and it does not score products. If a comment bank, an AES-LLM score, a dashboard, a “write better” chatbot, a self-applied rubric, or a marketing grader does not pass the four tests, it cannot be declared formative assessment.
7.1. Test of learning goals prior to use of the model, not of the comment bank or the demo. Balducci (2024) places student assessment in a human-centered design prior to the tool. Davies (2026) works on drafts: the draft belongs to a task someone defined. Pan (2025) indexes to a rubric: if it is prior, it can orient; if it is one more output, it cannot. Zhang et al. (2025) measure student AI literacy: an inventory does not replace the subject goal. The test is met when, before pasting the model onto the draft, it is written what is to be learned, which error is not delegated, and which uses are out of bounds, for example certifying. If the only criterion is a fluent comment or a demo that impressed, the test fails.
7.2. Test of interpretive criteria revisable by the teacher, not of the AES score or the prompting anchor. Pack et al. (2024), Li and Liu (2024), and Tate et al. (2024) speak of validity, reliability, or usefulness of a score: that is not the criterion of this assignment. Choi et al. (2026) make the anchor the key to accessible AES: the anchor is not discussed with the group. Kim and Holden (2026) change API and GUI: the interface does not interpret. Alfredo et al. (2024) ask for real participation in the design. The test is met when the teacher can read, amend, and teach what counts as performance in this genre, and a disagreement between score and human reading is not resolved in favor of the number by default. If they only see a holistic or a self-applied rubric, there is display scoring, not a criterion.
7.3. Test of dialogue and of the revision cycle with the learner, not of the chatbot that rewrites or the dashboard that colors. Guo et al. (2024) and Bauer et al. (2023) hold peer work as an exchange. Banihashem et al. (2024) compare sources: the source enters the cycle or stays in the margin. Xu et al. (2024) visualize: seeing is not conversing. Prilop and Weber (2023) train the comment with an expert. Davies (2026) requires, by the object of the draft itself, a next version. Goslen et al. (2025) generate plans: the plan is not the revision of the essay. The test is met when the learner receives the feedback, can ask, produces an attributable revision, and someone reads it. If the chatbot delivers the “better” paragraph and that version is published, there is substitution of the process. If the dashboard colors and no one speaks, there is similarity surveillance, not mediation.
7.4. Test of explicit attribution of evaluative judgment: who decides the grade or progress. Cui et al. (2026) oblige one to articulate collaborative competence: the framework does not offer a subject on whom to load the grade without naming them. Yan et al. (2024) do not offer an innocuous automated grade. Capel and Brereton (2023) do not dispense with the person. Kokoç et al. (2026) and Yang and Banks (2025) leave literacy with teachers and educators: the gap asks for formation, not cession of the name. Balducci (2024) keeps the evaluative act in an education of persons. The test is met if it is written who adopts, corrects, or discards the score and the comment, and when progress is declared —or denied— under a proper name. “The AI graded,” with no name and no threshold, is a disclaimer. The four are read together: a goal without a criterion does not apply; a criterion without dialogue is a file; dialogue without attribution does not hold the judgment tomorrow; attribution without a prior goal is invented under pressure.
8. Discussion
Three tensions organize the discussion. The first opposes displaying artifacts and exercising the craft of mediation. Pack et al. (2024), Li and Liu (2024), Choi et al. (2026), Kim and Holden (2026), and Tate et al. (2024) do not say the same thing: validity, reliability, anchor, interface, holistic usefulness. The market crushes them into “the grader already assesses.” The piece may be true in its domain; “piece = formative assessment” is not. Davies (2026) and Pan (2025) can show generative returns without authorizing that equation. Banihashem et al. (2024) can contrast sources without making AI the subject of the cycle.
The second opposes the prior layers of the series to formative mediation. The harness produces checkable work. Agent evaluation measures reliability on a distribution. HITL governs delegation in situation. All three can be well made and the feedback of this task can still fail to form: because the score was not the goal (Pack et al., 2024; Tate et al., 2024), because no one dialogued the comment (Guo et al., 2024; Prilop and Weber, 2023), because the rubric was self-applied (Pan, 2025), or because the name of who grades was left empty (Cui et al., 2026; Yan et al., 2024). Buying harness, eval, and HITL does not buy formative assessment. Omitting the fourth layer —“the model already comments, therefore it already forms”— is a display bias, not a conclusion of the reviews.
The third opposes evaluative literacy and prompt confidence. Kokoç et al. (2026) find an empirical field of in-service teacher literacy: it is a map, not a return protocol. Yang and Banks (2025) work literacy in elementary teacher educators as self-study: the coherent response is formation with a criterion, not a workshop of declared confidence. Zhang et al. (2025) validate a student inventory: measuring it is not teaching how to use feedback. Xu et al. (2024) recall that visualizing assessment data does not replace the craft. Alfredo et al. (2024), Capel and Brereton (2023), and Balducci (2024) keep human-centering as design and as act, not as a label. Bauer et al. (2023) and Goslen et al. (2025) offer neighbors —NLP for the peer, plans for scaffolding— that fit inside the craft and do not exhaust it. Confusing enthusiasm of use with mediation leaves the classroom governed by whoever most wants to delegate the judgment.
The tests read those tensions as hypotheses, not as a finding from a school. Choi et al. (2026) did not write a subject rubric. Kim and Holden (2026) did not study revision cycles. The peer work of Guo et al. (2024) is not the signature of a grade. Goal, criterion, dialogue, and name are a candidacy for formative assessment. Bank, score, dashboard, chatbot, or grader, alone, are consumption.
9. Limits
The review is narrative. It does not estimate combined effects and does not apply a PRISMA of its own. Sample sizes, d, r, AUC, agreement percentages, and improvement rates are not invented, even when some sources report them. Pack et al. (2024), Li and Liu (2024), Choi et al. (2026), Kim and Holden (2026), and Tate et al. (2024) are read by the object they score, not as trials of a formative protocol. Guo et al. (2024), Banihashem et al. (2024), Xu et al. (2024), and Prilop and Weber (2023) speak of peer work, of sources, or of training: effects are not imported. Zhang et al. (2025) validate a construct. Kokoç et al. (2026) map; they do not audit products. Yang and Banks (2025) are a self-study of elementary teacher educators, not a kindergarten study and not a trial of a grader.
Other sources are floor or neighbor, not direct evidence of the protocol. Davies (2026) and Pan (2025) describe uses of generative feedback: they do not certify the four tests. Bauer et al. (2023) propose a framework and an agenda. Cui et al. (2026) assess collaborative competence in Chinese EFL. Goslen et al. (2025) generate plans in games. Alfredo et al. (2024), Yan et al. (2024), Capel and Brereton (2023), and Balducci (2024) map human-centered design and challenges; they do not install a subject revision cycle. In this kit there is no trial that equates bank, score, dashboard, or chatbot with formative assessment under the four tests. That absence is a category ceiling, not a magnitude of harm. Contracts are not audited and brands are not condemned. The four tests are not a validated instrument. The “garden” label in heading 5 does not transfer early-childhood findings. The corpus is in English and the synthesis is editorial in Spanish. The stack of harness, agent eval, HITL, and formative mediation is a distinction of the series, not an empirical result of the twenty-one sources.
10. Conclusions
An automated comment bank, an LLM rubric score, a similarity dashboard, or a “write better” chatbot does not constitute formative assessment if teacher mediation is missing. Neither does a self-applied rubric nor a marketing grader. The craft is the judgment that decides which feedback matters, when to return it, and how to turn it into learning evidence. Pack et al. (2024), Li and Liu (2024), Choi et al. (2026), Kim and Holden (2026), Tate et al. (2024), Davies (2026), Pan (2025), and Banihashem et al. (2024) fix the ceiling: scoring, anchor, interface, holistic usefulness, generative return and rubric indexing, and sources that are not, by themselves, a cycle.
Where there is craft there is a subject and a protocol. Prilop and Weber (2023), Guo et al. (2024), Bauer et al. (2023), and Xu et al. (2024) leave the comment —expert, peer, visualized, or NLP-supported— as an object someone holds, not as a verdict of the model. Kokoç et al. (2026), Yang and Banks (2025), and Zhang et al. (2025) move literacy from the click to teachers, educators, and students, without equating it with evaluative literacy. Cui et al. (2026) oblige one to articulate collaboration. Alfredo et al. (2024), Capel and Brereton (2023), Balducci (2024), Yan et al. (2024), and Goslen et al. (2025) forbid closing the problem with a catalog role or a generated plan.
The stack is not substituted from within. The harness produces work. Agent evaluation measures reliability. HITL governs the copilot. Formative mediation turns feedback into learning evidence, with four conjunctive tests: goals prior to the model; interpretive criteria revisable by the teacher; dialogue and a revision cycle; attribution of who decides the grade or progress. Where the sources did not measure that protocol, this article does not take it as measured. Where they measured scores, sources, peer work, instruments, or frameworks, they are not translated into “the grader already forms.” Turning a generative system into formative assessment is exercising judgment about which comment is published, when it is revised, and who signs progress. The rest is a bank, a score, a board, or a marketing chatbot. It is not craft, and it must not be presented as what it is not.
Editorial Laboratory of NEXTECH.IA / Ingeniero Mitre.
References
- Alfredo, R., Echeverria, V., Jin, Y., Yan, L., Swiecki, Z., Gašević, D., y Martinez-Maldonado, R. (2024). Human-centred learning analytics and AI in education: A systematic literature review. Computers and Education: Artificial Intelligence, 6, Article 100215. https://doi.org/10.1016/j.caeai.2024.100215
- Balducci, B. (2024). AI and student assessment in human-centered education. Frontiers in Education, 9, Article 1383148. https://doi.org/10.3389/feduc.2024.1383148
- Banihashem, S. K., Kerman, N. T., Noroozi, O., Moon, J., y Drachsler, H. (2024). Feedback sources in essay writing: peer-generated or AI-generated feedback? International Journal of Educational Technology in Higher Education, 21(1), Article 23. https://doi.org/10.1186/s41239-024-00455-4
- Bauer, E., Greisel, M., Kuznetsov, I., Berndt, M., Kollar, I., Dresel, M., Fischer, M. R., y Fischer, F. (2023). Using natural language processing to support peer-feedback in the age of artificial intelligence: A cross-disciplinary framework and a research agenda. British Journal of Educational Technology, 54(5), 1222–1245. https://doi.org/10.1111/bjet.13336
- Capel, T., y Brereton, M. (2023). What is human-centered about human-centered AI? A map of the research landscape. En Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (pp. 1–23). ACM. https://doi.org/10.1145/3544548.3580959
- Choi, J., Tate, T., y Warschauer, M. (2026). Anchor is the key: Toward accessible automated essay scoring with large language model through prompting. Assessing Writing, 69, Article 101053. https://doi.org/10.1016/j.asw.2026.101053
- Cui, H., Mahfoodh, O. H. A., Dong, H., y Ulanbek, M. (2026). An assessment framework for human-AI collaborative competence in second language writing: Evidence from the Chinese EFL context. System, 137, Article 103937. https://doi.org/10.1016/j.system.2025.103937
- Davies, R. (2026). Using generative AI for formative feedback on student writing drafts. Educational Technology Research and Development. https://doi.org/10.1007/s11423-026-10652-9
- Goslen, A., Kim, Y. J., Rowe, J., y Lester, J. (2025). LLM-based student plan generation for adaptive scaffolding in game-based learning environments. International Journal of Artificial Intelligence in Education, 35(2), 533–558. https://doi.org/10.1007/s40593-024-00421-1
- Guo, K., Pan, M., Li, Y., y Lai, C. (2024). Effects of an AI-supported approach to peer feedback on university EFL students' feedback quality and writing ability. The Internet and Higher Education, 63, Article 100962. https://doi.org/10.1016/j.iheduc.2024.100962
- Kim, J., y Holden, D. (2026). Reassessing automated essay scoring with large language models: Evidence from API and GUI interfaces. Assessing Writing, 69, Article 101080. https://doi.org/10.1016/j.asw.2026.101080
- Kokoç, M., Ağırman, Ş., y Yağmurlugil, B. N. (2026). Mapping in-service teacher AI literacy: A systematic review of empirical studies. Teaching and Teacher Education, 181, Article 105707. https://doi.org/10.1016/j.tate.2026.105707
- Li, W., y Liu, H. (2024). Applying large language models for automated essay scoring for non-native Japanese. Humanities and Social Sciences Communications, 11(1), Article 723. https://doi.org/10.1057/s41599-024-03209-9
- Pack, A., Barrett, A., y Escalante, J. (2024). Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability. Computers and Education: Artificial Intelligence, 6, Article 100234. https://doi.org/10.1016/j.caeai.2024.100234
- Pan, Y. (2025). Leveraging generative AI powered rubric-indexed feedback as a formative assessment strategy for enhancing medical English education. Discover Computing, 28(1), Article 284. https://doi.org/10.1007/s10791-025-09830-9
- Prilop, C. N., y Weber, K. E. (2023). Digital video-based peer feedback training: The effect of expert feedback on pre-service teachers’ peer feedback beliefs and peer feedback quality. Teaching and Teacher Education, 127, Article 104099. https://doi.org/10.1016/j.tate.2023.104099
- Tate, T. P., Steiss, J., Bailey, D., Graham, S., Moon, Y., Ritchie, D., Tseng, W., y Warschauer, M. (2024). Can AI provide useful holistic essay scoring? Computers and Education: Artificial Intelligence, 7, Article 100255. https://doi.org/10.1016/j.caeai.2024.100255
- Xu, L., Zou, X., y Hou, Y. (2024). Effects of feedback visualisation of peer-assessment on pre-service teachers' data literacy, learning motivation, and cognitive load. Journal of Computer Assisted Learning, 40(4), 1447–1462. https://doi.org/10.1111/jcal.12955
- Yan, L., Sha, L., Zhao, L., Li, Y., Martinez-Maldonado, R., Chen, G., Li, X., Jin, Y., y Gašević, D. (2024). Practical and ethical challenges of large language models in education: A systematic scoping review. British Journal of Educational Technology, 55(1), 90–112. https://doi.org/10.1111/bjet.13370
- Yang, S., y Banks, A. (2025). Enhancing AI literacy: A collaborative self-study of elementary teacher educators. Studying Teacher Education, 21(3), 297–317. https://doi.org/10.1080/17425964.2024.2434213
- Zhang, H., Perry, A., y Lee, I. (2025). Developing and validating the Artificial Intelligence Literacy Concept Inventory: An instrument to assess artificial intelligence literacy among middle school students. International Journal of Artificial Intelligence in Education, 35(1), 398–438. https://doi.org/10.1007/s40593-024-00398-x