Summary: Discrete choice models
This post offers a 5-minute summary of the concepts I learned in the discrete choice analysis class by Dr. Bhat in Fall 2014 at the University of Texas at Austin.
Discrete choice models seek to predict choices made by individuals given the attributes of an individual (and the alternative). Consider an example: say, we want to predict which mode will a traveler with income level $x_1$ and age group $x_2$ choose, given the alternatives are drive alone (DA), car-pool (CP), or use transit (TR). The objective of DCMs is to find the parameters of a model given data on the choices of different individuals:
$y=\beta x+\epsilon$
where $y$ is a discrete variable indicating the selected choice. A regular regression model won't be a good choice to predict this as $y$ is discrete. I recently learned that the researchers in the machine learning field call this problem logistic regression, but we will stick with the econometrics term DCM (which, I believe, is applied more in the context when actual humans make decisions than robots!!).
We start with four elements in the choice process:
Discrete choice models seek to predict choices made by individuals given the attributes of an individual (and the alternative). Consider an example: say, we want to predict which mode will a traveler with income level $x_1$ and age group $x_2$ choose, given the alternatives are drive alone (DA), car-pool (CP), or use transit (TR). The objective of DCMs is to find the parameters of a model given data on the choices of different individuals:
$y=\beta x+\epsilon$
where $y$ is a discrete variable indicating the selected choice. A regular regression model won't be a good choice to predict this as $y$ is discrete. I recently learned that the researchers in the machine learning field call this problem logistic regression, but we will stick with the econometrics term DCM (which, I believe, is applied more in the context when actual humans make decisions than robots!!).
We start with four elements in the choice process:
- Decision maker: For the mode choice problem, this will be an individual. Though a decision maker can be a household, a firm, a government agency depending on the context.
- Alternatives: These are the available alternative to each decision maker. We define something called evoked choice set which includes only the alternatives that a traveler actually considers.
- Attributes: There are two types. Alternative specific attribute pertain to the alternative like travel time or cost of a transit bus or car-pool, and individual specific attributes pertain to the individual like their income, age etc.
- Decision rule: What rule does a decision maker employ to make the decision. This includes the rule of dominance (choose alternative $i$ if its utility is better than $j$) or the rule of satisfaction (choose the first alternative $i$ if it meets all my satisfaction criteria; example, selecting a 2BR house)
In the course, we focused on the utility-maximizing theory (based on the rule of dominance; call the theory UMT for now). One key idea in UMT is that only the relative magnitude of an alternative compared to the others matters; the numerical utility value has no tangible meaning. For example, car-pool will be preferred over driving alone if $U_{CP} = 50, U_{DA}=10$ and also if $U_{CP} = -100, U_{DA}=-150$. To explain the variation in choices made by individuals in the population, DCM defines the utility for an alternative perceived by an individual to have a deterministic component (based on the variables we can observe) and a stochastic component which varies from one individual to the other. In the equation on the very beginning $\beta x$ would be the deterministic component and $\epsilon$ the stochastic component.
Then, the rest of the course is Math! Here are some brief highlights:
- We first define the utility wrt each alternative
- While making an estimation of parameters of the utility of each alternative, we collect the parameters for alternative specific attributes together by subtracting the utility of one alternative from the other because those are assumed same for an individual. This simplifies the estimation process.
- The output of the model is $y$ which indicates the probability of choosing alternative $i$ whose attributes are supplied in the variable $x$.
- Say we have the mode choice data for $N$ individuals, we define a likelihood function, which gives an estimate of the likelihood of the observed data given the values of the parameters are fixed. The objective of maximum likelihood estimation is then to find the parameters which maximize this likelihood. A sample likelihood function looks like $L(\beta)= \Pi_{i=1}^N P_i(.)$, where P_i(.) is the probability that the individual $i$ will choose the alternative it actually chose in the data set.
- In simpler cases, when the stochastic term has a certain distribution, we can express the likelihood function analytically and differentiate it to find the right values of parameters. The famous logit model and probit model are derived by assuming error term to have Gumbel and Gaussian distribution, respectively.
- The goodness of the fit is measured by tracking the log-likelihood value for a given estimate and other statistics like adjusted $\rho_c^2$.
- Other advanced concepts from the course include: calculating elasticity terms (like the value of time); conducting hypothesis tests to find the significance of a certain parameter in the model; dealing with partial and full market segmentation; and defining advanced DCMs like nested logit and mixed logit based on relaxing the assumptions of the correlation of error terms and stochasticity of coefficients respectively.
The interesting questions that I seek to investigate after this revision are:
- The DCMs for mode choice are validated using stated or revealed preference surveys. However, for route choice, there is a lack of data availability which explains which individual chose which route. I wonder what other ways are used to train a model without directly using the individual choices, like possibly comparing link flows produced from the model with the field counts which is often done in traffic assignment.
- Dynamic discrete choice models are more fascinating because they involve solving a dynamic program in each step of parameter estimation. I wonder how complex is maximum likelihood estimation in those settings. The next step is to read this survey paper on DDCMs.
Comments
Post a Comment