数智时代的政策学习方法探索——以最低生活保障为例Policy Learning in the Digital Age: A Case Study of Precise Policy-Making in the China's Minimum Livelihood Guarantee
方悦,胡诗云,苏良军,解海天
摘要(Abstract):
政策学习作为人工智能技术与因果推断理论的深度融合,为大数据时代的精准社会治理提供了科学依据。本文立足于中国治理规模巨大、政策约束条件复杂的现实背景,系统梳理了政策学习领域的国际前沿文献,并提出一种基于凸化处理理论的政策学习算法。该方法有效解决了在复杂高维政策空间中的优化计算难题,显著提升了大数据及多约束场景下的政策学习效率。文章进一步以中国最低生活保障制度为例,利用中国家庭金融调查数据进行实证分析。结果表明,该方法能够生成具备高度可解释性的精准低保分配方案,有效扩大社会福利增益并提升治理效能。本研究为运用前沿数智技术推动中国社会治理现代化提供了重要参考。
关键词(KeyWords): 数字治理;政策学习;因果推断;大数据方法;最低生活保障
基金项目(Foundation): 国家自然科学基金青年项目“政策学习中的理论创新与实践应用”(72503208)、国家自然科学基金重点项目“高维计量模型的机器学习方法及其在经济管理中的应用”(72133002)、国家自然科学基金青年项目“基于机器学习的因果政策制定——政策学习”(72403008)和重大项目“大规模商务场景下的统计学习与管理实践”(72495123)的资助
作者(Author): 方悦,胡诗云,苏良军,解海天
参考文献(References):
- 陈强.2025.计量经济学中的因果推断:过去、现在与未来[J].中山大学学报(社会科学版),65(1):43-64.Chen Q.2025.Causal inference in econometrics:The past,the present and the future[J].Journal of Sun YatSen University(Social Science Edition),65(1):43-64.(in Chinese) 符华平,顾海.2009.我国农村最低生活保障制度的现状、问题及对策[J].南京社会科学,(1):101-104.Fu H P,Gu H.2009.Current situation,problems and countermeasures of China's rural minimum living security system[J].Social Sciences in Nanjing,(1):101-104.(in Chinese) 甘犁,尹志超,贾男,等.2013.中国家庭资产状况及住房需求分析[J].金融研究,(4):1-14.Gan L,Yin Z C,Jia N,et al.2013.Assets and residential demand of Chinese households[J].Journal of Financial Research,(4):1-14.(in Chinese) 韩华为.2021.代理家计调查、农村低保瞄准精度和减贫效应---基于中国家庭金融调查的实证研究[J].社会保障评论,5(2):93-109.Han H W.2021.Proxy means tests analysis of rural Dibao's targeting performance and anti-poverty effectiveness:empirical study based on China household finance survey[J].Chinese Social Security Review,5(2):93-109.(in Chinese) 胡诗云,江弘毅,解海天.2026.双重机器学习的理论与应用---从“黑箱”到“工具箱”的实践指南[J].数量经济技术经济研究,43(3):154-179.Hu S Y,Jiang H Y,Xie H T.2026.The theory and practices of double machine learning:From“black box”to“tool box”[J].Journal of Quantitative&Technological Economics,43(3):154-179.(in Chinese) 李棉管.2017.技术难题、政治过程与文化结果---“瞄准偏差”的三种研究视角及其对中国“精准扶贫”的启示[J].社会学研究,32(1):217-241.Li M G.2017.Technological barriers,political process and cultural consequence:Three research perspectives on targeting error and the implications for“targeting poverty”in China[J].Sociological Studies,32(1):217-241.(in Chinese) 民政部.2023.2023年民政事业发展统计公报[R/OL].https://www.mca.gov.cn/n156/n2679/c1662004999980001204/attr/355717.pdf.Ministry of Civil Affairs.2023.Statistical bulletin on the development of civil affairs in 2023[R/OL].https://www.mca.gov.cn/n156/n2679/c1662004999980001204/attr/355717.pdf.(in Chinese) 宋锦,李实,王德文.2020.中国城市低保制度的瞄准度分析[J].管理世界,36(6):37-48.Song J,Li S,Wang D W.2020.Analysis on the targeting of urban Dibao program in China[J].Journal of Management World,36(6):37-48.(in Chinese) 陶旭辉,郭峰.2023.异质性政策效应评估与机器学习方法:研究进展与未来方向[J].管理世界,39(11):216-235.Tao X H,Guo F.2023.Policy effect heterogeneity evaluation and machine learning methods:Research progress and future orientation[J].Journal of Management World,39(11):216-235.(in Chinese) 姚建平.2018.中国城市低保瞄准困境:资格障碍、技术难题,还是政治影响?[J].社会科学,(3):61-72.Yao J P.2018.The targeting dilemma of urban Dibao in China:Eligibility obstruction,technological problem,or political influences?[J].Journal of Social Sciences,(3):61-72.(in Chinese) 张征宇,陈浩文,郎旭华.2026.面向共同富裕问题的因果推断[J].经济评论,(1):3-19.Zhang Z Y,Chen H W,Lang X H.2026.Causal inference on the issue of common prosperity[J].Economic Review,(1):3-19.(in Chinese) 赵明华,田北海.2024.农村低保瞄准偏差估计及乡村治理手段创新的改进效应---基于双重机器学习的因果推断[J].华中农业大学学报(社会科学版),(5):167-180.Zhao M H,Tian B H.2024.Estimation of targeting bias for rural Dibao and the improvement effect of rural governance innovation:Causal inference based on double machine learning[J].Journal of Huazhong Agricultural University(Social Sciences Edition),(5):167-180.(in Chinese) 周文明,谢圣远.2016.中国城镇居民最低生活保障制度的发展演进及政策评估[J].广东社会科学,(2):206-212.Zhou W M,Xie S Y.2016.The development and policy evaluation of China's urban minimum living security system[J].Social Sciences in Guangdong,(2):206-212.(in Chinese) 朱梦冰,李实.2017.精准扶贫重在精准识别贫困人口---农村低保政策的瞄准效果分析[J].中国社会科学,(9):90-112.Zhu M B,Li S.2017.The key to precise poverty alleviation rests in the precise identification of impoverished populations:An analysis of the targeting effect of the rural subsistence allowance policy[J].Social Sciences in China,(9):90-112.(in Chinese) Ai C R,Fang Y,Xie H T.2026.Data-driven policy learning for continuous treatments[J].Journal of Econometrics,253:106170. Athey S,Wager S.2021.Policy learning with observational data[J].Econometrica,89(1):133-161. Chen Z,Jiang X,Liu Z K,et al.2023.Tax policy and lumpy investment behaviour:Evidence from China's VATreform[J].The Review of Economic Studies,90(2):634-674. DubéJ P,Misra S.2023.Personalized pricing and consumer welfare[J].Journal of Political Economy,131(1):131-189. Fang Y,Xi J,Xie H T.2025.Model selection for multivalued-treatment policy learning in observational studies[J].Journal of Business&Economic Statistics,43(4):897-909. Flores C A,Flores-Lagunes A,Gonzalez A,et al.2012.Estimating the effects of length of exposure to instruction in a training program:The case of job corps[J].Review of Economics and Statistics,94(1):153-171. Haushofer J,Niehaus P,Paramo C,et al.2025.Targeting impact versus deprivation[J].American Economic Review,115(6):1936-1974. Ida T,Ishihara T,Ito K,et al.2026.Choosing who chooses:Selection-driven targeting in energy rebate programs[J].Econometrica,94(1):225-247. Kaboski J P,Townsend R M.2011.A structural evaluation of a large-scale quasi-experimental microfinance initiative[J].Econometrica,79(5):1357-1406. Kallus N,Zhou A.2021.Minimax-optimal policy learning under unobserved confounding[J].Management Science,67(5):2870-2890. Kitagawa T,Tetenov A.2018.Who should be treated?Empirical welfare maximization methods for treatment choice[J].Econometrica,86(2):591-616. Kitagawa T,Tetenov A.2021.Equality-minded treatment choice[J].Journal of Business&Economic Statistics,39(2):561-574. Kitagawa T,Sakaguchi S,Tetenov A.2023.Constrained classification and policy learning[Z].ar Xiv preprint ar Xiv:2106.12886. Mbakop E,Tabord-Meehan M.2021.Model selection for treatment choice:Penalized welfare maximization[J].Econometrica,89(2):825-848. Tan-Soo J S,Cheng M D,Chen S.2025.Evaluating the labor supply implications of a cash transfer program:Evidence from China[J].Journal of Development Economics,176:103513. Zhan R H,Ren Z M,Athey S,et al.2024.Policy learning with adaptively collected data[J].Management Science,70(8):5270-5297. Zhou Z Y,Athey S,Wager S.2023.Offline multi-action policy learning:Generalization and optimization[J].Operations Research,71(1):148-183.
- ① 出自习近平总书记2024年4月22日至24日在重庆考察时的讲话,参见:https://www.gov.cn/yaowen/liebiao/202404/content_6947266.htm. (1)参见:https://www.ndrc.gov.cn/fggz/jyysr/jysrsbxf/202304/t20230428_1355158.html。