Showing 1–2 of 2 results for author: Meinfelder, F
-
On the Improvement of Predictive Modeling Using Bayesian Stacking and Posterior Predictive Checking
Authors:
Mariana Nold,
Florian Meinfelder,
David Kaplan
Abstract:
Model uncertainty is pervasive in real world analysis situations and is an often-neglected issue in applied statistics. However, standard approaches to the research process do not address the inherent uncertainty in model building and, thus, can lead to overconfident and misleading analysis interpretations. One strategy to incorporate more flexible models is to base inferences on predictive modeli…
▽ More
Model uncertainty is pervasive in real world analysis situations and is an often-neglected issue in applied statistics. However, standard approaches to the research process do not address the inherent uncertainty in model building and, thus, can lead to overconfident and misleading analysis interpretations. One strategy to incorporate more flexible models is to base inferences on predictive modeling. This approach provides an alternative to existing explanatory models, as inference is focused on the posterior predictive distribution of the response variable. Predictive modeling can advance explanatory ambitions in the social sciences and in addition enrich the understanding of social phenomena under investigation. Bayesian stacking is a methodological approach rooted in Bayesian predictive modeling. In this paper, we outline the method of Bayesian stacking but add to it the approach of posterior predictive checking (PPC) as a means of assessing the predictive quality of those elements of the stacking ensemble that are important to the research question. Thus, we introduce a viable workflow for incorporating PPC into predictive modeling using Bayesian stacking without presuming the existence of a true model. We apply these tools to the PISA 2018 data to investigate potential inequalities in reading competency with respect to gender and socio-economic background. Our empirical example serves as rough guideline for practitioners who want to implement the concepts of predictive modeling and model uncertainty in their work to similar research questions.
△ Less
Submitted 29 February, 2024;
originally announced February 2024.
-
Data Fusion for Joining Income and Consumption Information Using Different Donor-Recipient Distance Metrics
Authors:
Florian Meinfelder,
Jannik Schaller
Abstract:
Data fusion describes the method of combining data from (at least) two initially independent data sources to allow for joint analysis of variables which are not jointly observed. The fundamental idea is to base inference on identifying assumptions, and on common variables which provide information that is jointly observed in all the data sources. A popular class of methods dealing with this partic…
▽ More
Data fusion describes the method of combining data from (at least) two initially independent data sources to allow for joint analysis of variables which are not jointly observed. The fundamental idea is to base inference on identifying assumptions, and on common variables which provide information that is jointly observed in all the data sources. A popular class of methods dealing with this particular missing-data problem is based on nearest neighbour matching. However, exact matches become unlikely with increasing common information, and the specification of the distance function can influence the results of the data fusion. In this paper we compare two different approaches of nearest neighbour hot deck matching: One, Random Hot Deck, is a variant of the covariate-based matching methods which was proposed by Eurostat, and can be considered as a 'classical' statistical matching method, whereas the alternative approach is based on Predictive Mean Matching. We discuss results from a simulation study to investigate benefits and potential drawbacks of both variants, and our findings suggest that Predictive Mean Matching tends to outperform Random Hot Deck.
△ Less
Submitted 30 November, 2020;
originally announced December 2020.