Selection endogenous dummy ordered probit, and selection endogenous dummy dynamic ordered probit models

Similar documents
Endogenous Treatment Effects for Count Data Models with Endogenous Participation or Sample Selection

Endogenous Treatment Effects for Count Data Models with Sample Selection or Endogenous Participation

NELS 88. Latent Response Variable Formulation Versus Probability Curve Formulation

A double-hurdle count model for completed fertility data from the developing world

Applied Econometrics Lecture 1

Applied Health Economics (for B.Sc.)

Longitudinal Data Analysis Using Stata Paul D. Allison, Ph.D. Upcoming Seminar: May 18-19, 2017, Chicago, Illinois

Comments on: Panel Data Analysis Advantages and Challenges. Manuel Arellano CEMFI, Madrid November 2006

When Should We Use Linear Fixed Effects Regression Models for Causal Inference with Longitudinal Data?

FIML estimation of an endogenous switching model for count data

Women. Sheng-Kai Chang. Abstract. In this paper a computationally practical simulation estimator is proposed for the twotiered

Econometric Analysis of Cross Section and Panel Data

DEEP, University of Lausanne Lectures on Econometric Analysis of Count Data Pravin K. Trivedi May 2005

Least Squares Estimation of a Panel Data Model with Multifactor Error Structure and Endogenous Covariates

Even Simpler Standard Errors for Two-Stage Optimization Estimators: Mata Implementation via the DERIV Command

Non-linear panel data modeling

David M. Drukker Italian Stata Users Group meeting 15 November Executive Director of Econometrics Stata

Ch 7: Dummy (binary, indicator) variables

Introduction to GSEM in Stata

Generalized Linear Latent and Mixed Models with Composite Links and Exploded

AGEC 661 Note Fourteen

Short T Panels - Review

Longitudinal Data Analysis Using SAS Paul D. Allison, Ph.D. Upcoming Seminar: October 13-14, 2017, Boston, Massachusetts

Estimating and Using Propensity Score in Presence of Missing Background Data. An Application to Assess the Impact of Childbearing on Wellbeing

Sensitivity checks for the local average treatment effect

Child Maturation, Time-Invariant, and Time-Varying Inputs: their Relative Importance in Production of Child Human Capital

A Course in Applied Econometrics Lecture 14: Control Functions and Related Methods. Jeff Wooldridge IRP Lectures, UW Madison, August 2008

Discrete Choice Modeling

When Should We Use Linear Fixed Effects Regression Models for Causal Inference with Longitudinal Data?

Discrete Choice Modeling

Adding Uncertainty to a Roy Economy with Two Sectors

A simple bivariate count data regression model. Abstract

Gibbs Sampling in Latent Variable Models #1

Simplified Implementation of the Heckman Estimator of the Dynamic Probit Model and a Comparison with Alternative Estimators

Multivariate and Multivariable Regression. Stella Babalola Johns Hopkins University

Estimating the Dynamic Effects of a Job Training Program with M. Program with Multiple Alternatives

ECON 594: Lecture #6

Generalized Linear Models for Non-Normal Data

Identification of Regression Models with Misclassified and Endogenous Binary Regressor

Recent Advances in the Field of Trade Theory and Policy Analysis Using Micro-Level Data

Logistic regression: Why we often can do what we think we can do. Maarten Buis 19 th UK Stata Users Group meeting, 10 Sept. 2015

leebounds: Lee s (2009) treatment effects bounds for non-random sample selection for Stata

Flexible Estimation of Treatment Effect Parameters

1 Motivation for Instrumental Variable (IV) Regression

Econometrics. Week 8. Fall Institute of Economic Studies Faculty of Social Sciences Charles University in Prague

-redprob- A Stata program for the Heckman estimator of the random effects dynamic probit model

Applied Microeconometrics (L5): Panel Data-Basics

Microeconometrics. Bernd Süssmuth. IEW Institute for Empirical Research in Economics. University of Leipzig. April 4, 2011

Econometrics of Panel Data

New Developments in Nonresponse Adjustment Methods

Dynamic analysis of binary longitudinal data

Limited Dependent Variables and Panel Data

Groupe de lecture. Instrumental Variables Estimates of the Effect of Subsidized Training on the Quantiles of Trainee Earnings. Abadie, Angrist, Imbens

An Introduction to Mplus and Path Analysis

Internal vs. external validity. External validity. This section is based on Stock and Watson s Chapter 9.

Introduction to Regression Analysis. Dr. Devlina Chatterjee 11 th August, 2017

Beyond the Target Customer: Social Effects of CRM Campaigns

Tobit and Selection Models

Microeconometrics. C. Hsiao (2014), Analysis of Panel Data, 3rd edition. Cambridge, University Press.

ECONOMETFUCS FIELD EXAM Michigan State University May 11, 2007

INTRODUCTION TO MULTILEVEL MODELLING FOR REPEATED MEASURES DATA. Belfast 9 th June to 10 th June, 2011

Econometrics Homework 4 Solutions

Causal Inference Lecture Notes: Causal Inference with Repeated Measures in Observational Studies

Categorical and Zero Inflated Growth Models

Tables and Figures. This draft, July 2, 2007

Write your identification number on each paper and cover sheet (the number stated in the upper right hand corner on your exam cover).

Introduction to Econometrics

Generated Covariates in Nonparametric Estimation: A Short Review.

Harvard University. A Note on the Control Function Approach with an Instrumental Variable and a Binary Outcome. Eric Tchetgen Tchetgen

Econometrics of Panel Data

Joint Modeling of Longitudinal Item Response Data and Survival

Treatment Effects with Normal Disturbances in sampleselection Package

Extended regression models using Stata 15

Path Analysis. PRE 906: Structural Equation Modeling Lecture #5 February 18, PRE 906, SEM: Lecture 5 - Path Analysis

Short Questions (Do two out of three) 15 points each

Empirical approaches in public economics

Wooldridge, Introductory Econometrics, 4th ed. Chapter 15: Instrumental variables and two stage least squares

Basic econometrics. Tutorial 3. Dipl.Kfm. Johannes Metzler

STA 216, GLM, Lecture 16. October 29, 2007

Estimation of Dynamic Nonlinear Random E ects Models with Unbalanced Panels.

Econometric Analysis of Games 1

Identification and Estimation of Nonlinear Dynamic Panel Data. Models with Unobserved Covariates

Chapter 1 Introduction. What are longitudinal and panel data? Benefits and drawbacks of longitudinal data Longitudinal data models Historical notes

Syllabus. By Joan Llull. Microeconometrics. IDEA PhD Program. Fall Chapter 1: Introduction and a Brief Review of Relevant Tools

Model Assumptions; Predicting Heterogeneity of Variance

Control Function and Related Methods: Nonlinear Models

Instrumental Variables

Outline. Overview of Issues. Spatial Regression. Luc Anselin

Lecture 11/12. Roy Model, MTE, Structural Estimation

Systematic error, of course, can produce either an upward or downward bias.

Semiparametric Estimation of a Sample Selection Model in the Presence of Endogeneity

Sociology 593 Exam 2 March 28, 2002

Citation for published version (APA): Jak, S. (2013). Cluster bias: Testing measurement invariance in multilevel data

ECO375 Tutorial 4 Wooldridge: Chapter 6 and 7

Introduction to causal identification. Nidhiya Menon IGC Summer School, New Delhi, July 2015

Econometrics in a nutshell: Variation and Identification Linear Regression Model in STATA. Research Methods. Carlos Noton.

Interpreting and using heterogeneous choice & generalized ordered logit models

REGRESSION RECAP. Josh Angrist. MIT (Fall 2014)

A Bayesian Nonparametric Approach to Monotone Missing Data in Longitudinal Studies with Informative Missingness

Appendix A: The time series behavior of employment growth

Transcription:

Selection endogenous dummy ordered probit, and selection endogenous dummy dynamic ordered probit models Massimiliano Bratti & Alfonso Miranda

In many fields of applied work researchers need to model an ordinal dependent variable that is only observed for a proportion of the sample and that is a function of an endogenous variable Smoking and Education, but not everybody smokes Drinking and Education, but not everybody drinks Job type (unskilled/skilled/proffesional) and Education, but not everybody works Selection can have different sources: Entry into an activity (smoking / drinking) Survey and/or item non-response Fundamentally, the DGP of selection into missingness and the DGP of the ordinal variable are essentially two different although related processes (e.g., extensive vs. intensive margin decisions)

Previous work Terza, Kenkel, Tsui-Fang and Shinichi (2008) suggest a twostep method for estimating a selection endogenous dummy model for an interval coded dependent variable (grouped). This is an extension of Mullahy (1998) Modified Two Part Model and Mullahy (1986) Hurdle Model. Despite being a two-step approach, this is not a LIML estimator but relies on joint multivariate normality. Miranda and Rabe-Hesketh (2006) consider a model for an ordinal dependent variable with either an endogenous dummy or sample selection. Cannot deal with the two problems at the same time. Harris and Zhao (2007) suggests a zero-inflated ordered probit model, which is quite similar to Mullahy s Hurdle model and is related to Lambert (1992) zero-inflated count data models. Bratti and Miranda (2009) use the BCS70 and the methods described here to analyse how higher education affects smoking intensity in the UK.

Selection endogenous dummy ordered probit Let y i be the ordinal variable of interest. Variable y i is generated according to a continuous latent variable model y i = x i β + δg i + v i, (1) where the observed response y i is determined by a threshold model missing if S i = 0 1 if y it k 1 & S i = 1 2 if k y i = 1 < yit k 2 & S i = 1...... H if k H 1 < yit & S i = 1, G i represents a potentially endogenous dummy and the main response y i is only observed if a selection rule S i = 1 is met. Both G i and S i are always observed and {k 1,, k H 1 } R H 1 are constants to be estimated along other parameters.

Selection endogenous dummy ordered probit II The endogenous dummy and the selection dummy are also generated according to a continuous latent variable model Si = r i θ + ϕg i + q i G i = z i γ + w i, with S i = 1(Si > 0) and G i = 1(Gi > 0). Although the model is identified by functional form it is valuable to specify a set of exclusion restrictions to avoid problems of tenuous identification. Hence, when possible, some elements of the z i should not enter x i or r i, and some elements of r i should not enter x i or z i. (2) Correlation among the three outcomes is allowed by imposing some structure to the error terms, w i = u i + ζ i, q i = λ 1 u i + η i v i = λ 2 u i + ν i, (3)

Selection endogenous dummy ordered probit III To close the model we require the covariates to be all strictly exogenous, and the idiosyncratic errors to be orthogonal given the individual heterogeneity term u, D(u x, z, r) = D(u) (4) D(ζ x, z, r, u) = D(ζ u) (5) D(η x, z, r, u) = D(η u) (6) D(ν x, z, r, u) = D(ν u) (7) ζ u η u ν u. (8) To ease estimation we suppose that u, ζ, η, and ν are all independent standard normal, in which case D(ζ u) = D(ζ), D(η u) = D(η), and D(ν u) = D(ν). The factors loadings {λ 1 λ 2 } R 2 are free parameters and allow any type of correlation (positive, negative or null) between yi, Gi and Si.

Selection endogenous dummy ordered probit IV Correlation between unobservables entering Gi and Si is given by: ρ gs = λ 1 2 ( ). 1 + λ 2 1 Similarly, correlation between unobservables entering G and is given by: y i ρ gy = λ 2 2 ( ). 1 + λ 2 2 Finally, correlation between unobservables entering main response y and selection Si is given by: ρ sy = λ 1 λ 2 (1 + λ 2 1 ) ( 1 + λ 2 2 ). In this model G is exogenous wrt S if ρ gs = 0, G is exogenous wrt y if ρ gy = 0, and y is observed at random if ρ sy = 0.

Selection endogenous dummy ordered probit V Estimate the system by Maximum Simulated Likelihood Analytical first derivatives and numerical second derivatives Can also do OPG approx. of the Hessian (much faster!) Halton sequences cover the (0,1) interval better and require fewer draws to achieve high precision than random samples from uniform distribution Program written in Stata/Mata Really fast! Stata 10/SE + 400 Halton draws + 2,792 indv / 8,043 pers-obs + numerical 2nd derivatiives = 1.6hrs Stata 10/SE + 400 Halton draws + 2,792 indv / 8,043 pers-obs + OPG Hessian = less than 5min

Selection Endogenous dummy dynamic ordered probit model Suppose that y i is observed for two periods t = {1, 2}. We now extend the model to accommodate the fact that the outcome of the ordinal response in period 2 can be a function of the value that the variable took in period 1. In other words, we consider the possibility of having autoregressive dynamics in the ordered variable y i2 such that y i1 = m i δ + δ 1G i + v i (9) H yi2 = x i β + δ 2G i + π j 1(y i1 = j) + ξ i, (10) j=1

Selection endogenous dummy dynamic ordered probit model II As usual we suppose the model is complemented by a threshold rule, missing if S it = 0 1 if y it k 1 & S it = 1 2 if k y it = 1 < yit k 2 & S it = 1...... H if k H 1 < yit & S it = 1, To ease presentation we suppose that G and S do not have dynamics themselves. Further, we suppose that y is always observed in the first period and that G is a time invariant variable. Hence, The selection and endogenous dummies are generated by the following latent variable models, Git = Gi = z i γ + w i (11) Sit = Si = r i θ + ϕ 1G i + q i (12)

Selection endogenous dummy dynamic ordered probit model III Like in the model, we impose some structure to the error terms, w i = u i + ζ i, q i = λ 1 u i + η i (13) v i = λ 2 u i + ν i ξ i = λ 3 u i + ε i, where u i, ζ i, η i, ν i, and ε i are all supposed to be independent standard normal, and λ 1,..., λ 3 are free factor loadings. Notice that the model fully recognises the fact that the initial state dummies1(y i1 = 1),..., 1(y i1 = H) are potentially endogenous in equation (10) by modelling together the initial conditions and the dynamic equation, and allowing any type of correlation among unobservables entering both equations. There is true state dependance if the coefficients on the initial state dummies in (10) are different from zero.

Selection endogenous dummy dynamic ordered probit model IV Unobservables entering the system are now correlated in six different ways: ρ g,s = λ 1 q 2(1+λ 2 1) ρ g,y1 = λ 2 q 2(1+λ 2 2) ρ g,y2 = λ 3 q 2(1+λ 2 3) ρ s,y1 = λ 1 λ 2 q(1+λ 2 1 )(1+λ 2 2) ρ s,y2 = λ 1 λ 3 q(1+λ 2 1 )(1+λ 2 3) ρ y1,y 2 = λ 2 λ 3 q(1+λ 2 2 )(1+λ 2 3).

Selection endogenous dummy dynamic ordered probit model V The initial state dummies can be included into the S without further complications The model can be extended to allow dynamics in S and G The model can deal with more than two periods with relative minor modifications As before we use MSL for estimation Program written in Stata/Mata Really fast!

I - Variables definition We apply the model to study the effect of higher education (HE) on drinking frequency using the British Cohort Study 1970 (BCS70), 29-year follow-up survey. Variables definition: y i (drinking frequency) = 1 (2 to 3 times a month); 2 (once a week); 3 (2 to 3 days a week); 4 (on most days) G i (higher education) = 1 (HE); 0 (lower than HE) S i (usual drinker) = 1 (drinks more than 2 to 3 times a month); 0 (drinks less than 2 to 3 times a month, i.e. less often or in special occasions, not nowadays, never drunk)

II - Descriptive statistics Descriptive statistics, BCS70, 29-year follow-up survey % of usual drinkers Women Men No HE 69.9 85.78 HE 83.56 92.02 Drinking frequency per month - usual drinkers 2-3 times once a 2-3 times per month week a week most days Total Women no HE 25.04 34.4 31.63 8.92 100 HE 16.01 24.53 43.53 15.93 100 Total 21.7 30.75 36.04 11.51 100 Men no HE 14.14 24.95 43.27 17.63 100 HE 8.95 19.13 48.35 23.57 100 Total 12.35 22.94 45.03 19.69 100

III - specification HE: parents absence, parents education, highest social class, school type, state, BAS score, ethnicity, religion, height, teacher s assessment of child s knowledge and parental interest in child s education, teacher s homework style, all at age 10. Usual drinker: HE, parents absence, parents education, highest social class, school type, state, ethnicity, religion, height, all at age 10; height at age 30, mother s drinking during pregnancy, homework on parents demand, month of 29-year follow-up interview Drinking frequency: same controls as the selection equation (usual drinker) Since the selection and the main outcome variables refer to the same process (drinking) we preferred not to impose exclusion restrictions between the two equations. In other cases, for instance item non-reponse and panel attrition, it can be easier to find valid exclusion restrictions.

III - Results Be a Drinking frequency -y- (ME) usual drinker -s- 2-3 times once a 2-3 times most (ME) a month week a week days Men HE -g- 0.156-0.129-0.109 0.054 0.183 [0.015] [0.032] [0.021] [0.009] [0.046] ρ gs -0.484 [0.056] ρ gy -0.320 [0.091] ρ sy 0.310 [0.063] No. obs. 3300 Women HE -g- 0.259-0.302-0.087 0.181 0.208 [0.024] [0.046] [0.010] [0.015] [0.039] ρ gs -0.418 [0.061] ρ gy -0.479 [0.087] ρ sy 0.401 [0.058] No. obs. 3525 ( ) Significant at 1% (5%). Eicker-Huber-White robust standard errors reported in square brackets. Marginal effects (ME) are computed at the sample mean. For dummy variables, they show the change in the relevant probability when the variable changes from 0 to 1.

Hence, our empirical application shows that: HE is endogenous wrt both drinking participation (i.e. be an usual drinker) and drinking frequency Unobservables affecting drinking participation and drinking frequency are positively correlated (positive selection) HE has a positive causal effect on both drinking participation and drinking frequency (i.e., educated people drink more) The effect is much larger for females than for males

References Bratti, M and Miranda, A (2009) Non-pecuniary returns to higher education: The effect on smoking intensity in the UK. Health Economics Published Online: Jul 13 2009 12:08PM (Early View) http://dx.doi. org/10.1002/hec.1529. Harris, MN and Zhao, X (2007) A zero-inflated ordered probit model, with an application to modelling tobacco consumption. Journal of Econometrics 141:1073 1099. Heckman, JJ (1979) Sample selection bias as a specication error. Econometrica 47: 153 61. Lambert,D (1992) Zero-Inflated Poisson Regression with an Application to Defects in Manufacturing. Technometrics 34:1-14. Miranda, A. and Rabe-Hesketh, S. (2006) Maximum likelihood estimation of endogenous switching and sample selection models for binary, ordinal, and count variables. Stata Journal 6 (3): 285 308. Terza, JV and Kenkel, DS and Lin, TF and Sakata, S (2008) Care-giver advice as a preventive measure for drinking during pregnancy: zeros, categorical outcome responses, and endogeneity. Health Economics 17: 41-54. Train, KE (2003). Discrete choice methods with simulation. Cambridge university press.