2026

LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models
LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models

Yida Cai, Xin Dai, Bingxiang He, Huiyuan Xie, Yuxiao Ye, Zhenghao Liu, Yang Bai, Zhiyuan Liu

Under review.

We introduce LexReward, a taxonomy-driven framework for legal reward modeling that evaluates response quality across three complementary dimensions: Style, Element, and Chain. We develop dimension-specific rubrics and use the resulting rewards to construct pairwise preference data for DPO and to train a family of reward models, LexRM. Experiments show that the rubric-based rewards reliably distinguish response quality and that DPO improves performance across all three dimensions. LexRM further enables reinforcement learning to improve policy performance in each corresponding dimension without requiring reference answers at reward time.

LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models

Yida Cai, Xin Dai, Bingxiang He, Huiyuan Xie, Yuxiao Ye, Zhenghao Liu, Yang Bai, Zhiyuan Liu

Under review.

We introduce LexReward, a taxonomy-driven framework for legal reward modeling that evaluates response quality across three complementary dimensions: Style, Element, and Chain. We develop dimension-specific rubrics and use the resulting rewards to construct pairwise preference data for DPO and to train a family of reward models, LexRM. Experiments show that the rubric-based rewards reliably distinguish response quality and that DPO improves performance across all three dimensions. LexRM further enables reinforcement learning to improve policy performance in each corresponding dimension without requiring reference answers at reward time.

LexIssue: Benchmarking Legal Issue Identification in Chinese Civil Litigation
LexIssue: Benchmarking Legal Issue Identification in Chinese Civil Litigation

Huiyuan Xie*#, Yuqin Huang*, Zhicheng Hao, Yida Cai, Shaochun Wang, Zhenghao Liu, Yuxiao Ye (* equal contribution, # corresponding author)

Under review.

We construct LexIssue, a benchmark containing 430 real-world Chinese civil litigation cases and 1,303 expert-annotated disputed legal issues. We further develop an issue-centric legal knowledge base spanning 27 causes of action and 441 candidate legal issue entries to support retrieval-augmented reasoning. Experimental results across a diverse set of models show that retrieval-augmented generation using the constructed legal issue knowledge base consistently improves performance in identifying disputed legal issues and their corresponding legal attributes.

LexIssue: Benchmarking Legal Issue Identification in Chinese Civil Litigation

Huiyuan Xie*#, Yuqin Huang*, Zhicheng Hao, Yida Cai, Shaochun Wang, Zhenghao Liu, Yuxiao Ye (* equal contribution, # corresponding author)

Under review.

We construct LexIssue, a benchmark containing 430 real-world Chinese civil litigation cases and 1,303 expert-annotated disputed legal issues. We further develop an issue-centric legal knowledge base spanning 27 causes of action and 441 candidate legal issue entries to support retrieval-augmented reasoning. Experimental results across a diverse set of models show that retrieval-augmented generation using the constructed legal issue knowledge base consistently improves performance in identifying disputed legal issues and their corresponding legal attributes.

LexRel: Benchmarking Legal Relation Extraction for Chinese Civil Cases
LexRel: Benchmarking Legal Relation Extraction for Chinese Civil Cases

Yida Cai, Ranjuexiao Hu, Huiyuan Xie#, Chenyang Li, Yun Liu, Yuxiao Ye, Zhenghao Liu, Weixing Shen, Zhiyuan Liu# (# corresponding author)

ACL2026 CCF-A

In this work, we firstly introduce a comprehensive schema, which contains a hierarchical taxonomy and definitions of arguments, for AI systems to capture legal relations in Chinese civil cases. Based on this schema, we formulate a legal relation extraction task and present LexRel, an expertannotated benchmark for legal relation extraction in the Chinese civil law domain.

LexRel: Benchmarking Legal Relation Extraction for Chinese Civil Cases

Yida Cai, Ranjuexiao Hu, Huiyuan Xie#, Chenyang Li, Yun Liu, Yuxiao Ye, Zhenghao Liu, Weixing Shen, Zhiyuan Liu# (# corresponding author)

ACL2026 CCF-A

In this work, we firstly introduce a comprehensive schema, which contains a hierarchical taxonomy and definitions of arguments, for AI systems to capture legal relations in Chinese civil cases. Based on this schema, we formulate a legal relation extraction task and present LexRel, an expertannotated benchmark for legal relation extraction in the Chinese civil law domain.

2025

Rank-then-score: Enhancing large language models for automated essay scoring
Rank-then-score: Enhancing large language models for automated essay scoring

Yida Cai, Kun Liang, Sanwoo Lee, Qinghan Wang, Yunfang Wu

Preprint

In this paper, we propose Rank-Then-Score (RTS), a fine-tuning framework based on large language models to enhance their essay scoring capabilities.

Rank-then-score: Enhancing large language models for automated essay scoring

Yida Cai, Kun Liang, Sanwoo Lee, Qinghan Wang, Yunfang Wu

Preprint

In this paper, we propose Rank-Then-Score (RTS), a fine-tuning framework based on large language models to enhance their essay scoring capabilities.

2024

Unleashing large language models’ proficiency in zero-shot essay scoring
Unleashing large language models’ proficiency in zero-shot essay scoring

Sanwoo Lee*, Yida Cai*, Desong Meng, Ziyang Wang, Yunfang Wu# (* equal contribution, # corresponding author)

EMNLP2024 CCF-B

In this paper, we show that our zero-shot prompting framework, Multi Trait Specialization (MTS), elicits LLMs’ ample potential for essay scoring.

Unleashing large language models’ proficiency in zero-shot essay scoring

Sanwoo Lee*, Yida Cai*, Desong Meng, Ziyang Wang, Yunfang Wu# (* equal contribution, # corresponding author)

EMNLP2024 CCF-B

In this paper, we show that our zero-shot prompting framework, Multi Trait Specialization (MTS), elicits LLMs’ ample potential for essay scoring.

Assessing the performance of Chinese open source large Language models in information extraction tasks

Yida Cai, Hao Sun, Hsiu-Yuan Huang, Yunfang Wu

Preprint.

Information Extraction (IE) plays a crucial role in Natural Language Processing (NLP) by extracting structured information from unstructured text, thereby facilitating seamless integration with various real-world applications that rely on structured data. Despite its significance, recent experiments focusing on English IE tasks have shed light on the challenges faced by Large Language Models (LLMs) in achieving optimal performance, particularly in sub-tasks like Named Entity Recognition (NER). In this paper, we delve into a comprehensive investigation of the performance of mainstream Chinese open-source LLMs in tackling IE tasks, specifically under zero-shot conditions where the models are not fine-tuned for specific tasks. Additionally, we present the outcomes of several few-shot experiments to further gauge the capability of these models. Moreover, our study includes a comparative analysis between these open-source LLMs and ChatGPT, a widely recognized language model, on IE performance. Through meticulous experimentation and analysis, we aim to provide insights into the strengths, limitations, and potential enhancements of existing Chinese open-source LLMs in the domain of Information Extraction within the context of NLP.

Assessing the performance of Chinese open source large Language models in information extraction tasks

Yida Cai, Hao Sun, Hsiu-Yuan Huang, Yunfang Wu

Preprint.

Information Extraction (IE) plays a crucial role in Natural Language Processing (NLP) by extracting structured information from unstructured text, thereby facilitating seamless integration with various real-world applications that rely on structured data. Despite its significance, recent experiments focusing on English IE tasks have shed light on the challenges faced by Large Language Models (LLMs) in achieving optimal performance, particularly in sub-tasks like Named Entity Recognition (NER). In this paper, we delve into a comprehensive investigation of the performance of mainstream Chinese open-source LLMs in tackling IE tasks, specifically under zero-shot conditions where the models are not fine-tuned for specific tasks. Additionally, we present the outcomes of several few-shot experiments to further gauge the capability of these models. Moreover, our study includes a comparative analysis between these open-source LLMs and ChatGPT, a widely recognized language model, on IE performance. Through meticulous experimentation and analysis, we aim to provide insights into the strengths, limitations, and potential enhancements of existing Chinese open-source LLMs in the domain of Information Extraction within the context of NLP.

2023

Task-Related Pretraining with Whole Word Masking for Chinese Coherence Evaluation

Ziyang Wang, Sanwoo Lee, Yida Cai, Yunfang wu

NLPCC

This paper presents an approach for evaluating coherence in Chinese middle school student essays, addressing the challenges of time-consuming and inconsistent essay assessment.

Task-Related Pretraining with Whole Word Masking for Chinese Coherence Evaluation

Ziyang Wang, Sanwoo Lee, Yida Cai, Yunfang wu

NLPCC

This paper presents an approach for evaluating coherence in Chinese middle school student essays, addressing the challenges of time-consuming and inconsistent essay assessment.