01
AI3L 2026 · APSCE TBICS · Kumamoto, Japan · 27 June 2026

Structured AI-Guided Essay Revision
with Dual-Rubric Assessment

A Human–AI Collaborative Approach to Evaluating Writing and Interaction Quality in an EFL Course

Simon Wang
Language Centre, Hong Kong Baptist University
simonwang@hkbu.edu.hk
Liping Deng (Lisa)
Dept of Education Studies, Hong Kong Baptist University
lisadeng@hkbu.edu.hk
📍 Room 1 (9F), Kumamoto University, Japan 📅 Saturday 27 June · 10:45–11:00 🏛️ APSCE TBICS · AI3L 7 · Paper #217S 🔗 Conference Programme
AI3L 7 Session · Chair: Yi-ju Ariel WU · Room 1 (9F) · 10:15–11:15
10:15–10:30 139S — Orchestrating Multimodal GenAI in Preservice English Teacher Education (Yi-ju Ariel WU)
10:30–10:45 216S — Narrative Scaffolding and Linguistic Expression in Students with ASD through AI-Generated Comics (Shou-Nan TAN et al.)
10:45–11:00 217S — Structured AI-Guided Essay Revision with Dual-Rubric Assessment (Simon WANG & Liping DENG) ← This talk
11:00–11:15 219S — AI Character-Embedded Metaverse for Nursing Education (Chia-Yi CHENG et al.)

Scan Slides

eegc.hkbu.me/27June.html

Try the Platform

eegc.hkbu.me/aiedit
02
Why We Built This

AI Has Enormous Potential — But Teachers Are Left Out of the Loop

The Opportunity

AI can engage every student in a genuine, personalised writing conversation — at scale.

For EFL learners who rarely get individual feedback, this has enormous potential.

⚠️

The Risk

Unconstrained AI use leads to overreliance — students accept suggestions passively rather than engaging critically.

The process of revision disappears.

🔑

The Core Gap

Teachers are left out of the loop.

No visibility into how students interact with AI. No structured workflow. No way to monitor or assess the process.

So how can teachers stay in the loop? → Design the workflow · Monitor the interaction · Build the tools yourself
03
Our Response

The Answer: Build Your Own Platform

Generic AI chat tools (ChatGPT, etc.) give students a private conversation the teacher can never see. The solution is not to ban AI — it is to design the environment.
🎛️

Control the Interaction

We write the system prompt — the hidden instruction that shapes everything the AI says. The AI follows our pedagogical logic, not the student's whims.

Three structured steps. No free-form rewriting. AI guides; student thinks.

🔍

Own the Data

Every message, every exchange, every score is saved to our database. Teachers see the full chat record — not just the final essay.

Nothing happens off-record. The process is as visible as the product.

⚙️

Do More with Data

With full access to chat + essay data, we can: use AI to score essays, measure interaction quality, flag low-engagement students, and conduct research — all from one system.

Data that would otherwise disappear becomes a research asset.

This is the design principle behind everything that follows — not a chatbot wrapper, but a purpose-built pedagogical platform where teachers stay in the loop at every step.
03
The Course

LANG 0036 — Remedial English at HKBU

Who Are These Students?

  • University freshmen at Hong Kong Baptist University who scored lower on English placement tests
  • Required to take LANG 0036 before the standard first-year Academic English course
  • Many have limited experience with argumentative or academic writing
  • EFL learners who rarely receive individual, sustained feedback on their writing

What Is the Course?

  • LANG 0036 — Enhancing English through Global Citizenship
  • Focus: academic writing, essay structure, argumentation, and critical thinking
  • 840 students enrolled across 48 sections in Semester 1, 2025–26
  • Taught by a large team — individual writing coaching at scale is impossible
The challenge: 840 students who need personalised writing feedback — in a course where teachers cannot give individual attention to every draft.
04
The Assessment Task

A Point-of-View Essay — Graded in Two Dimensions

20%
of final course grade
Point-of-view essay task
10%
Essay Quality
How much did the essay improve?
10%
AI Interaction Quality
How well did they engage with AI?
Rubric 1 — Essay Quality
Applied to original AND revised drafts
Content & Ideas
Organisation & Logical Progression
Vocabulary
Grammar & Sentence Structure
Scale: 1 (Limited) → 5 (Excellent) per dimension
Rubric 2 — AI Interaction Quality
Applied to the full chat history with AI
Conversation Depth
Critical Review of AI Suggestions
Refining Process
Scale: 1 (Limited) → 5 (Excellent) per dimension
05
The Platform — Live Demo

The EEGC AI Edit Module (Enhancing English through Global Citizenship)

Student
logs in
📋 Briefing
Learn the task
🔧 Training
Sample essay
📝 Assessment
Own essay
📄 AI Report
emailed
👩‍🏫 Teacher
Dashboard
01
1

Thesis Revision

AI evaluates the thesis for clarity and argumentative stance. Student must revise before proceeding — no skipping.

02
2

Topic Sentence

AI checks whether the topic sentence supports the thesis and is logically linked to the argument.

03
3

Paragraph Revision

Full revision: evidence, logic, vocabulary, grammar. AI gives feedback but does not rewrite.

👉 Open on your phone now
eegc.hkbu.me/aiedit
Tap 🎓 AI3L audience test to try it as a student — free AI tutoring, no sign-up
What the speaker will do now
  • Open the platform live
  • Log in with the audience guest token
  • Start Training Mode and show the AI interaction
06
Deployment at Scale — Semester 1, 2025–26

Who Used It?

840
Students enrolled
48
Active sections
734
Reports generated
(some students completed both modes)
Assessment Mode Completion 377 / 840  (44.9%)
Training Mode Completion 131 / 840  (15.6%)

Report Contents

  • Complete chat history (student ↔ AI)
  • AI-generated scores on both rubrics
  • AI contribution analysis (qualitative)
  • Original and revised essay drafts

Delivery

  • Emailed to student's HKBU address
  • PDF + Markdown format
  • Teacher dashboard: section-level view
  • 1,162 files total across 734 sessions
07
Preliminary Finding

Does Interaction Quality Predict Improvement?

The Pattern

Students with higher interaction rubric scores (Conversation Depth, Critical Review, Refining Process) also demonstrated greater essay quality improvement between original and revised drafts.

Why This Matters

  • Justifies Rubric 2 as a diagnostic tool, not just a participation grade
  • Teachers get a leading indicator: low interaction → likely weak revision
  • Supports structured revision as a replicable pedagogical model

⚠️ Important Caveat

This is correlational, not causal. Stronger writers may both engage more deeply with AI and show greater improvement.

Strongest predictor — Refining Process
r = 0.48 n = 101 assessed sessions
What is Refining Process? It's one of three dimensions in Rubric 2 (AI Interaction Quality). It measures whether students went back to the AI multiple times to refine their draft — not just accepting the first suggestion, but iterating: asking follow-up questions, trying alternative phrasings, and pushing for better outcomes.
What r = 0.48 means: Among 101 student sessions, those who scored higher on Refining Process also showed significantly larger essay score gains — even after we removed the effect of how strong their writing was to begin with. A moderate-to-strong positive link: more iteration with the AI → more essay improvement.
All three Rubric 2 dimensions vs essay improvement:   Refining Process r = 0.48  ·  Critical Review r = 0.41  ·  In-Depth Conversation r = 0.33
09
Framework

A Replicable Human–AI Collaboration Model

🤖 AI Handles

  • Scalable formative feedback (840 students, instant)
  • Structured three-step revision guidance
  • First-pass rubric scoring (both rubrics)
  • Report generation and delivery
  • Interaction quality logging

👩‍🏫 Teachers Handle

  • Quality assurance on AI scores
  • Section-level progress monitoring
  • Targeted intervention for struggling students
  • Final grade decisions
  • Pedagogical design of the module
Student submits essay
AI guides 3-step revision
Dual rubric scores generated
Report emailed
Teacher reviews flagged cases
"AI handles scalable formative feedback and initial scoring while teachers focus on quality assurance and pedagogical intervention." — This is the replicable model.
10
Takeaways

Implications & Next Steps

For Practice

  • Three-step mandatory sequencing prevents surface-only revision
  • Dual rubric opens a window into how students learn with AI
  • Model is replicable: any EFL/EAP writing course with LMS integration
  • Interaction rubric score can flag disengaged students before final grading

For Research

  • Interaction analytics as a new lens on AI-mediated learning
  • Kappa results point to where AI needs human calibration
  • Partial ρ suggests interaction quality has independent predictive value

Limitations

  • Correlation ≠ causation: interaction ↔ improvement
  • Vocabulary/Grammar kappa needs further calibration
  • One semester, one institution, one task type
  • Privacy: reports contain full chat histories — institutional consent obtained

Next Steps

  • Sem 2 data collection (ongoing)
  • Deeper chat history pattern analysis
  • Longitudinal writing improvement tracking
  • Refining interaction rubric operationalisation
10
How We Built It

Vibe-Coded — A One-Stop Platform for Engagement & Research

💬

Engage

Students interact with the AI tutor directly in their browser — no app install, no account setup beyond a student ID.

Three structured modes: Briefing → Training → Assessment

🗄️

Collect

Every session is auto-saved to a database and the full report is emailed to the student and teacher the moment it's generated.

Zero manual data collection. 734 reports. One semester.

📊

Analyze

Teacher dashboard + backend analytics. Query scores, correlations, and interaction patterns — all from the same system that ran the course.

The data on this slide was pulled live from that database.

Built entirely with AI-assisted (vibe) coding — from prompt design to rubric engine to report delivery — by a team of language teachers, not software engineers.
eegc.hkbu.me  ·  Nuxt 4 + Poe API + PostgreSQL · open for replication
12
Proposal · Honest Trade-offs

A Platform Any Teacher Can Use — What It Would Take

🏗️ What We Can Build

An online form-based platform where teachers configure an agentic tutor without writing a single line of code.

 Fill in: task type · rubric · topic · feedback style
 Platform generates a shareable student URL
 Students open link → interact with AI tutor in browser

No app install. No account for students. Works on any device.

Honest Constraints
🔑

BYOK — Bring Your Own Key

Teachers or students must provide their own API token (OpenAI, Gemini, Claude…). The platform runs it client-side — we never see or bill for AI usage.

Trade-off: low barrier to deploy, but students need an account with an AI provider.

📧

Email as Storage

No database. At session end, the full chat history is emailed to the student and teacher. Data stays in inboxes — not on our servers.

Trade-off: GDPR-friendly and zero hosting cost, but no aggregate analytics.

Open Questions

Is BYOK a dealbreaker?
Students at well-funded institutions may have free access via Copilot or Gemini. Others may not — does this reproduce the digital divide?

What do teachers lose without storage?
No cross-class trends. No rubric correlation analysis like we ran here. Email works for individual feedback — not for research.

Middle ground?
Teacher pays a small subscription → gets managed storage + shared API quota. Students never need their own key.

Ref: OpenAI for Education · "Hello Education, Meet Agents" & "How to Build Agents for Higher Education"
13
Under the Hood · Actual System Prompts

How We Told the AI to Scaffold, Not Ghost-write

Training Mode — Anti-ghostwriting Rule
If a student asks you to "help me revise", "rewrite this", "fix my essay"
→ Politely decline. Explain that your role is to guide them so they build the skills themselves.

Offer 2–3 short contrasting examples to illustrate direction so they can see improvement without you writing theirs.
Assessment Mode — Student-led Priority
Before starting revision:
① Negotiate targets — ask the student about their personal goals.
② Diagnose — review against rubric, identify strengths & gaps.
③ Let student choose — which weaknesses to focus on.

Never provide a fully rewritten paragraph or sentence.
Rubric B — Critical Review of AI Suggestions
5 All AI suggestions thoroughly evaluated; strong justification for acceptance and rejection 3 Some evaluated; partial justification 1 All suggestions accepted blindly — no critical thought
Accepting every AI suggestion earns the lowest score. Disagreeing thoughtfully earns the highest.
The Design Principle
The AI was contractually prevented from writing for the student. The rubric was designed to reward pushback. Together, they force the student to stay in the driver's seat.
Source: promptAndEssay.js — Training_Mode_Prompt, Assessment_Mode_Prompt, AssessBot_Prompt
14
AI3L 2026 · Paper #217

Thank You

Structured AI-Guided Essay Revision with Dual-Rubric Assessment: A Human–AI Collaborative Approach to Evaluating Writing and Interaction Quality in an EFL Course

Simon Wang
Language Centre, Hong Kong Baptist University
simonwang@hkbu.edu.hk
Liping Deng (Lisa)
Dept of Education Studies, Hong Kong Baptist University
lisadeng@hkbu.edu.hk
🌐 smartlessons.hkbu.tech/LANG0036/ 📊 840 students · 48 sections · 734 reports
AI-guided revision automated essay scoring human–AI collaboration formative feedback EFL writing dual-rubric assessment
Questions welcome
✏️ Inline Edit Mode — click any text to edit
🔏 Slide Editor

Enter the editor password to unlock slide editing.

Enable inline editing to click directly on any text in the slides and type. When done, save to persist changes.


Save captures the current slide deck content (including your edits) and writes it to disk. The page will reload to reflect the saved state.

Not authenticated.

📱 Scan to view slides

📱

Please Rotate Your Device

These slides are designed for landscape mode. Turn your phone sideways for the best experience.