isr
5/27
Debiasing ML- or AI-Generated Regressors in Partially Linear Models
텐센트, 아마존 웹서비스, 마이크로소프트 실험 시스템의 부분선형 회귀모형에서 자동라벨과 사람 주석 자료를 활용했다. 인공지능 예측을 그대로 쓰면 측정오류로 인과효과를 과대 또는 과소 추정하지만, 새 방법은 이를 바로잡았다. 대규모 자동라벨의 확장성을 유지하면서 신뢰할 인과추정과 인공지능 공정성 개선을 가능하게 한다.
Abstract
Organizations increasingly use machine learning (ML) and artificial intelligence (AI), including large language models, to generate variables for regression models that inform business and policy decisions. For example, practitioners may use AI to predict review sentiment, ad aesthetics, or emotional expressions, and then estimate their causal effects on outcomes such as sales or engagement. However, because AI predictions are imperfect, directly using these AI-generated variables as regressors introduces measurement error that can systematically bias causal estimates, potentially leading to over- or underinvestment in business strategies. We develop new estimators that correct this bias in partially linear regression models, which are widely deployed in experimental systems at major platforms, including Tencent, Amazon AWS, and Microsoft. Our approach requires only a small human-annotated subsample alongside the large AI-labeled data set to achieve unbiased and efficient estimation. We demonstrate that our methods work with both traditional ML algorithms and LLM-based predictions. Our framework can be directly integrated into existing analytics and experimental systems, enabling practitioners to leverage the scalability of AI-generated data while maintaining reliable causal conclusions. This work also has implications for AI fairness, as our approach can help correct biases from any source in AI predictions.
isr
5/28
When Behavioral Data Betray Users: A Diagnostic and Protective Framework Against Social Interaction Leakages
공개 행동자료에서 숨은 사회적 연결을 찾아내고 보호하는 틀을 미국 루이지애나와 펜실베이니아의 옐프 평점·후기 자료로 검증한다. 잘못 연결한 비율이 10%일 때 연결의 약 절반, 20%일 때 60% 이상을 찾아내며, 표적형 피싱 수익을 109%에서 1,098%로 높인다. 개인정보를 수학적으로 보호하는 두 방식은 자료 활용성을 유지하면서 연결 유출과 공격 수익을 낮춰 공유 위험 평가를 가능하게 한다.
Abstract
Digital platforms often share publicly visible behavioral data, such as ratings and reviews, to improve service quality. Yet, these seemingly innocuous data can also reveal hidden social ties among users. We develop a two-stage framework that first diagnoses this leakage risk by inferring latent social interactions from behavioral data and then mitigates the risk while preserving data utility. We characterize when hidden ties are identifiable from observed actions and show, using public Yelp data from Louisiana and Pennsylvania, that an attacker can recover about half of true social ties at a 10% false-positive rate and more than 60% at a 20% false-positive rate. These inferred ties can materially increase cyber risk: when used for spear phishing, the estimated return on attack rises from 109% for a campaign with 500 impersonation attempts to 1,098% for one with 10,000 attempts. To reduce this leakage, we propose two perturbation mechanisms with formal differential privacy guarantees on the released representations: one adds Gaussian noise directly to the released action matrix, and the other adds Laplace noise to the learned representation before resynthesizing the released data. Both mechanisms reduce the link-inference accuracy and substantially lower estimated attacker returns; under our main protection settings, returns become negative for smaller campaigns and remain much lower at larger scales. Overall, the paper provides a practical framework for diagnosing social interaction leakage and evaluating the privacy-utility trade-off when platforms share behavioral data. History: Peiyu Chen, Senior Editor; Heng Xu, Associate Editor. Funding: Y. Leng is supported by the U.S. National Science Foundation (NSF) [Grant IIS-2153468]. Supplemental Material: The online appendix is available at https://doi.org/10.1287/isre.2024.1469 .
isr
5/28
Enhancing the Benefits of Similarity for Marginalized Identities: How Technology Features Influence Online Reviews When One Evaluates a Similar Other
에어비앤비 숙박 후기를 분석해 게스트와 호스트의 주변화된 정체성 공유와 플랫폼 설계를 살폈다. 정체성 유사성만으로는 긍정적 후기가 늘지 않았지만 유연한 취소와 즉시 예약은 이를 높였다. 플랫폼은 대표성보다 거래의 통제권과 유연성을 기술 구조에 반영해 포용성을 높여야 한다.
Abstract
This study examines how similarity, social identity, and platform design jointly shape user experiences in two-sided digital platforms. Drawing on theories of power asymmetry, social identity, and digital platforms, we analyze more than 500,000 reviews on Airbnb to evaluate how shared marginalized identities between guests and hosts influence perceived experiences. Contrary to expectations, similarity alone does not consistently yield more positive outcomes; however, platform features that enable power sharing, such as flexible cancellation policies and instant booking, significantly increase the likelihood of positive reviews among marginalized groups. These findings contain practical and policy implications. For platform designers and policymakers, the results highlight the opportunity of structural features that redistribute control and flexibility within transactions to foster inclusive and positive experiences. For platform owners and regulators, these results suggest a shift from representation-focused interventions toward feature-based governance that embeds power sharing, and potentially other shared values, into platform architecture. More broadly, the study underscores the importance of aligning technological features with the values and lived experiences of marginalized users, offering a scalable pathway to promote both inclusion and marketplace performance.
isr
5/28
An Alternative Model of Brand Loyalty
브랜드 충성도를 전환비용으로 생긴 충성도와 반복 사용으로 제품을 선호하게 된 충성도로 나누어 시장 결과를 비교한다. 전환비용은 기존 기업을 보호해 경쟁을 약화하고 가격을 높여 후생을 줄이지만, 선호 변화는 차별화를 줄여 경쟁과 후생을 높이고 가격을 낮출 수 있다. 따라서 기업의 사용자 경험·신뢰성·제품 설계 투자는 잠금 효과 투자와 다르며, 규제는 모든 고객 고착이 아닌 전환 장벽을 겨냥해야 한다.
Abstract
There are two different types of brand loyalty, and they work in different ways. Our study shows that loyalty created by switching costs, such as proprietary standards, data frictions, or restrictive interfaces, has very different market consequences from loyalty created by preference shifts, where repeated use makes customers genuinely prefer a product’s design or features. Although both increase retention, switching costs protect incumbents, soften competition, raise prices, and reduce consumer and social welfare. Preference shifts can have the opposite effect; they can lead to a reduction in differentiation between rival products, intensify competition, lower prices, and under some conditions, improve welfare by better aligning products with what customers want. The managerial message is clear; investments in user experience, reliability, and product design are not strategically equivalent to investments in lock-in. The policy message is equally important; regulators should not treat all customer stickiness as harmful. Remedies, such as interoperability, data portability, or limits on proprietary standards, are best targeted at switching barriers, not loyalty created by genuine product improvement.