호텔 고객 리뷰와 기존 리뷰 답변 기록을 바탕으로, 사람의 선호에 맞는 대규모 언어 모델 미세조정법을 개발했다. 문맥 보강은 사실 오류를 줄였고, 새 학습 설계와 제약 방식은 기존 방법보다 나은 답변을 만들었다. 고객 응답을 빠르고 일관되게 지원할 가능성을 보이지만, 안전장치와 사람의 감독, 분야별 모델 조정이 필요하다.
Abstract
Online reviews can shape where people stay, eat, and shop, but businesses often struggle to keep up with the flood of customer feedback. Although generative artificial intelligence (AI) offers a promising solution, general-purpose models are not designed for the specific judgment, tone, and accuracy required in customer review responses. This study introduces a new fine-tuning method that helps large language models generate review replies that better match human preferences in real business settings. The paper makes several technical advances. It identifies why review-response systems hallucinate and introduces a context-augmentation strategy to reduce factual errors. It also develops a theory-driven way to automatically construct preference data from existing review-response records, overcoming a major barrier in preference fine-tuning. In addition, the study proposes a curriculum learning design and a new support-constraint method that reduces the overconservatism of existing offline optimization approaches, with stronger theoretical guarantees. Tests on hotel reviews show that the method produces better responses than leading alternatives in both automated evaluations and human judgments. The findings point to a practical path for using AI to help firms respond faster and more consistently to customers while also underscoring the need for safeguards, human oversight, and domain-specific model alignment in customer-facing AI systems.