
Yida Cai, Xin Dai, Bingxiang He, Huiyuan Xie, Yuxiao Ye, Zhenghao Liu, Yang Bai, Zhiyuan Liu
Under review.
We introduce LexReward, a taxonomy-driven framework for legal reward modeling that evaluates response quality across three complementary dimensions: Style, Element, and Chain. We develop dimension-specific rubrics and use the resulting rewards to construct pairwise preference data for DPO and to train a family of reward models, LexRM. Experiments show that the rubric-based rewards reliably distinguish response quality and that DPO improves performance across all three dimensions. LexRM further enables reinforcement learning to improve policy performance in each corresponding dimension without requiring reference answers at reward time.
Yida Cai, Xin Dai, Bingxiang He, Huiyuan Xie, Yuxiao Ye, Zhenghao Liu, Yang Bai, Zhiyuan Liu
Under review.
We introduce LexReward, a taxonomy-driven framework for legal reward modeling that evaluates response quality across three complementary dimensions: Style, Element, and Chain. We develop dimension-specific rubrics and use the resulting rewards to construct pairwise preference data for DPO and to train a family of reward models, LexRM. Experiments show that the rubric-based rewards reliably distinguish response quality and that DPO improves performance across all three dimensions. LexRM further enables reinforcement learning to improve policy performance in each corresponding dimension without requiring reference answers at reward time.

Huiyuan Xie*#, Yuqin Huang*, Zhicheng Hao, Yida Cai, Shaochun Wang, Zhenghao Liu, Yuxiao Ye (* equal contribution, # corresponding author)
Under review.
We construct LexIssue, a benchmark containing 430 real-world Chinese civil litigation cases and 1,303 expert-annotated disputed legal issues. We further develop an issue-centric legal knowledge base spanning 27 causes of action and 441 candidate legal issue entries to support retrieval-augmented reasoning. Experimental results across a diverse set of models show that retrieval-augmented generation using the constructed legal issue knowledge base consistently improves performance in identifying disputed legal issues and their corresponding legal attributes.
Huiyuan Xie*#, Yuqin Huang*, Zhicheng Hao, Yida Cai, Shaochun Wang, Zhenghao Liu, Yuxiao Ye (* equal contribution, # corresponding author)
Under review.
We construct LexIssue, a benchmark containing 430 real-world Chinese civil litigation cases and 1,303 expert-annotated disputed legal issues. We further develop an issue-centric legal knowledge base spanning 27 causes of action and 441 candidate legal issue entries to support retrieval-augmented reasoning. Experimental results across a diverse set of models show that retrieval-augmented generation using the constructed legal issue knowledge base consistently improves performance in identifying disputed legal issues and their corresponding legal attributes.

Yida Cai, Ranjuexiao Hu, Huiyuan Xie#, Chenyang Li, Yun Liu, Yuxiao Ye, Zhenghao Liu, Weixing Shen, Zhiyuan Liu# (# corresponding author)
ACL2026 CCF-A
In this work, we firstly introduce a comprehensive schema, which contains a hierarchical taxonomy and definitions of arguments, for AI systems to capture legal relations in Chinese civil cases. Based on this schema, we formulate a legal relation extraction task and present LexRel, an expertannotated benchmark for legal relation extraction in the Chinese civil law domain.
Yida Cai, Ranjuexiao Hu, Huiyuan Xie#, Chenyang Li, Yun Liu, Yuxiao Ye, Zhenghao Liu, Weixing Shen, Zhiyuan Liu# (# corresponding author)
ACL2026 CCF-A
In this work, we firstly introduce a comprehensive schema, which contains a hierarchical taxonomy and definitions of arguments, for AI systems to capture legal relations in Chinese civil cases. Based on this schema, we formulate a legal relation extraction task and present LexRel, an expertannotated benchmark for legal relation extraction in the Chinese civil law domain.

Yida Cai, Kun Liang, Sanwoo Lee, Qinghan Wang, Yunfang Wu
Preprint
In this paper, we propose Rank-Then-Score (RTS), a fine-tuning framework based on large language models to enhance their essay scoring capabilities.
Yida Cai, Kun Liang, Sanwoo Lee, Qinghan Wang, Yunfang Wu
Preprint
In this paper, we propose Rank-Then-Score (RTS), a fine-tuning framework based on large language models to enhance their essay scoring capabilities.

Sanwoo Lee*, Yida Cai*, Desong Meng, Ziyang Wang, Yunfang Wu# (* equal contribution, # corresponding author)
EMNLP2024 CCF-B
In this paper, we show that our zero-shot prompting framework, Multi Trait Specialization (MTS), elicits LLMs’ ample potential for essay scoring.
Sanwoo Lee*, Yida Cai*, Desong Meng, Ziyang Wang, Yunfang Wu# (* equal contribution, # corresponding author)
EMNLP2024 CCF-B
In this paper, we show that our zero-shot prompting framework, Multi Trait Specialization (MTS), elicits LLMs’ ample potential for essay scoring.
Yida Cai, Hao Sun, Hsiu-Yuan Huang, Yunfang Wu
Preprint.
Information Extraction (IE) plays a crucial role in Natural Language Processing (NLP) by extracting structured information from unstructured text, thereby facilitating seamless integration with various real-world applications that rely on structured data. Despite its significance, recent experiments focusing on English IE tasks have shed light on the challenges faced by Large Language Models (LLMs) in achieving optimal performance, particularly in sub-tasks like Named Entity Recognition (NER). In this paper, we delve into a comprehensive investigation of the performance of mainstream Chinese open-source LLMs in tackling IE tasks, specifically under zero-shot conditions where the models are not fine-tuned for specific tasks. Additionally, we present the outcomes of several few-shot experiments to further gauge the capability of these models. Moreover, our study includes a comparative analysis between these open-source LLMs and ChatGPT, a widely recognized language model, on IE performance. Through meticulous experimentation and analysis, we aim to provide insights into the strengths, limitations, and potential enhancements of existing Chinese open-source LLMs in the domain of Information Extraction within the context of NLP.
Yida Cai, Hao Sun, Hsiu-Yuan Huang, Yunfang Wu
Preprint.
Information Extraction (IE) plays a crucial role in Natural Language Processing (NLP) by extracting structured information from unstructured text, thereby facilitating seamless integration with various real-world applications that rely on structured data. Despite its significance, recent experiments focusing on English IE tasks have shed light on the challenges faced by Large Language Models (LLMs) in achieving optimal performance, particularly in sub-tasks like Named Entity Recognition (NER). In this paper, we delve into a comprehensive investigation of the performance of mainstream Chinese open-source LLMs in tackling IE tasks, specifically under zero-shot conditions where the models are not fine-tuned for specific tasks. Additionally, we present the outcomes of several few-shot experiments to further gauge the capability of these models. Moreover, our study includes a comparative analysis between these open-source LLMs and ChatGPT, a widely recognized language model, on IE performance. Through meticulous experimentation and analysis, we aim to provide insights into the strengths, limitations, and potential enhancements of existing Chinese open-source LLMs in the domain of Information Extraction within the context of NLP.
Ziyang Wang, Sanwoo Lee, Yida Cai, Yunfang wu
NLPCC
This paper presents an approach for evaluating coherence in Chinese middle school student essays, addressing the challenges of time-consuming and inconsistent essay assessment.
Ziyang Wang, Sanwoo Lee, Yida Cai, Yunfang wu
NLPCC
This paper presents an approach for evaluating coherence in Chinese middle school student essays, addressing the challenges of time-consuming and inconsistent essay assessment.