Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Saturday, July 9, 2011

Creating hysteria with fake statistics

The Village Voice goes after Ashton Kutcher, Demi Moore, their celebrity charity consultants, a host of credulous big-name media organizations and a bunch of faux charities enriching themselves with government grants for promulgating essentially meaningless statistics regarding child prostitution in the US.

This sort of statistical foolishness is hardly new nor is it limited to attacks on the sex industry; indeed, it accompanies most every media scare, from upcoming ice ages to dangerous vegetables to Satanism.

Ashton Kutcher - I'd never heard of him but then I mostly stopped watching TV other than sports and elections in the early 1980s - responds to the Village Voice piece like a spoiled child rather than like, to pick a phrase, a real man. Maybe he needs a twitter consultant / editor too.

A better reaction would be for Ashton and Demi to apologize to the public and then to give some money to the American Statistical Association to fund their educational efforts. With a large enough donation, perhaps the ASA would even set up a special program called "The Ashton and Demi Program in Basic Statistics for Celebrities and Journalists."

Wednesday, June 15, 2011

Genmatch

Genetic Matching for Estimating Causal Effects:
A General Multivariate Matching Method for Achieving Balance in
Observational Studies
Alexis Diamond
ŽJasjeet S. Sekhon

Abstract

This paper presents Genetic Matching, a method of multivariate matching, that uses an evolutionary search algorithm to determine the weight each covariate is given. Both propensity score matching and matching based on Mahalanobis distance are limiting cases of this method. The algorithm makes transparent certain issues that all matching methods must confront. We present simulation studies that show that the algorithm improves covariate balance, and that it may reduce conditional bias if the selection on observables assumption holds. We then present a reanalysis of a number of datasets in the LaLonde (1986) controversy.

JEL classification: C13, C14, H31

Keywords: Matching, Propensity Score, SeKeywords: Matching, Propensity Score, Selection on Observables, Genetic Optimization, Causal Inference
I happened to read this paper (for the second time) a couple of days ago. It introduces for an economist audience (the authors are political scientists) a new algorithm for the construction of estimates of causal effects based on an assumption of "selection on observed variables". Other disciplines sometimes call this assumption unconfoundedness or ignorability (and economists sometimes refer to the "conditional independence assumption"). What all this jargon means is that the researcher thinks that conditional on covariates available in the data, individuals are assigned to treatment, whether by nature, by institutions, or by their own choices, or some combination of these, in a way that is unrelated to their untreated outcomes. Essentially, one makes the case for an assumption random assignment conditional on observed characteristics, where in good papers that case consists of more than "this is all I could do" or "look at how many different variables I matched on mom!".

To see what makes the method outlined in this paper (and in more technical detail in other papers available on Jas Sekhon's webpage, including some co-authored with my UM political science colleague Walter Mebane) it helps to think about what a randomized experiment does. Statistically, random assignment of individuals into treatment and control groups balances the distributions of both observed and unobserved covariates between the treated and control units. When samples are small, this balance may be imperfect in particular realizations, but as the sample gets larger, the balance, statistically speaking, becomes better and better.

What Genmatch does is to choose untreated units to match to the treated units in an observational study based solely on a criterion of post-match balance. This contrasts with the usual approach in economics of using something like nearest neighbor matching or kernel matching on estimated propensity scores (probabilities of treatment) in an iterative process in which balance is tested at each iteration and the propensity score model is made more flexible by adding additional terms until some desired level of balance is achieved.

The paper includes two Monte Carlo analyses as well as an application to the much-abused data from LaLonde's (1986) seminal work on the National Supported Work Demonstration. The NSW application is well done and sensibly interpreted. One of the Monte Carlo analyses, drawn from the literature on matching outside of economics, has the bizarre feature that it includes matching on instruments, that is, on variables that affect participation in treatment but do not otherwise affect outcomes. As Jay Bhattacharya explains at length, you should not do this.

Readers interested in Genmatch will also likely be interested in inverse probability tilting, which has the same spirit of building balance maximization into the estimation but in the context of weighting estimators rather than matching estimators.

Saturday, May 21, 2011

Atlantic Causal Inference Conference

I spent the last two days at the Atlantic Causal Inference conference which, this year, was held far away from the Atlantic at Michigan's School of Public Health.

The conference was great inter-disciplinary fun, as it included statisticians, biotstatisticians, epidemiologists and even a few economists. I got to meet some famous people from other fields and also learned a lot (though there were a couple of presentations where simultaneous inter-disciplinary translation would have helped).

Two quotes that I liked well enough to write down:

"My colleagues, they study artificial intelligence; me, I study natural stupidity" - Amos Tversky

"Matching is an attempt to approximate what reweighting is doing directly" - Justin McCrary

In addition to providing the Tversky quote, Sander Greenland's keynote speech taught me the term "nullism", which indicates excessive reverence for the null hypothesis in classical hypothesis testing. I expect that term to come in handy in future.

Monday, May 16, 2011

Thursday, April 7, 2011

Multiple comparisons

The multiple comparisons problem arises when doing large numbers of statistical tests. For example, one might estimate the impact of some treatment on 100 different outcomes. Ignoring correlations among the outcomes, in a world where the population treatment effect equals zero for all 100 outcomes, one would still expect five estimates to be statistically significant at the five percent level.


My friend (and fellow Heckman student) Peter Schochet of Mathematica prepared a very nice (and very accessible) survey of the related literature for the Institute of Education Sciences.

Wednesday, March 16, 2011

Meaningless statistic increases

Someday perhaps someone will explain to me why anyone should pay attention to life expectancy numbers based on synthetic cohort assumptions. Under the synthetic cohort assumption, it is assumed that the mortality behavior of people born today can be well approximated by the mortality behavior of people who are old today. This is ludicrous. It ignores all technical change in health care as well as the major behavioral changes that we know exist across cohorts due to, for example, changes in BMI and in smoking. How can these numbers do other than mislead?

Thursday, February 24, 2011

Steve Stigler on robustness

A fine, short meditation on the history of robustness in statistics, motivated by recent work in economics.

Steve Stigler is the son of Nobel economist George Stigler, whose courses I took (and graded for) when I was a gradual student at Chicago. Steve shares his father's writing skills and his interest in intellectual history.

I took (well, audited) Steve's course on (surprise!) robust estimation shortly after I started working with Jim Heckman. It was in that course that I met statistics gradual student Nancy Clements, who ended up, at my instigation, becoming an RA for Heckman and a co-author on the 1997 Review of Economic Studies paper on heterogeneous treatment effects.

A couple of years after auditing his class, I read Steve Stigler's excellent book on The History of Statistics, which I continue to recommend to students at the start of every new econometrics course, whether graduate or undergraduate, that I teach. The book influenced the way I motivate least squares regression, which he frames in the book as a solution to the historical problem of what to do when you want to estimate a line but have more than the two data points required for identification.

Wednesday, January 26, 2011

200 years of progress in a few minutes

A very cool video from BBC 4 on what has happened to the relationship between life expectancy and income over the past 200 years.

One can, of course quibble about the different scales on the horizontal and vertical axis, and with the failure to note the lagging progress of sub-Saharan Africa amidst all the optimism at the end.

And then life expectancy is itself a very odd thing, based as it is on strong synthetic cohort assumptions.

But worth watching in any case, just for the graphical presentation.

Hat tip: Ken Troske

Thursday, December 9, 2010

Yankees hat = crime?

Or could it just be that the denominator has been mislaid?

The article manages to actually hint at the denominator in one paragraph, but it seems to me that it is the whole story.

Non-Yankees hat tip to Charlie Brown

Monday, December 6, 2010

Gelman on regression practice

A fine post from Andrew Gelman on important rules for regression practice.

You'll see some comments from ECONJEFF as well.

Thursday, November 4, 2010

Collaborating with bio-statisticians



Actually, the most challenging part of the "giving statistical / econometric advice" role is when it becomes clear to you as the advice-giver that in fact the study you are providing advice about is a waste of time. That is a hard message to pass along in a nice way.

Hat tip: Charlie Brown

Saturday, October 30, 2010

On matching

Chris Blattman has a fine rant about matching as a statistical method for program evaluation. That in turn engendered a response from Andrew Gelman and a reply from Blattman. This post is my response to them both and to the broader questions they raise in their post.

Perhaps the nice thing about all this is that the three of us generally agree on the main point: matching is not a magic bullet. Just because you estimate a propensity score and then run psmatch2 in Stata does not make the selection on observed variables assumption any more true than it was when you were running a linear regression of the outcome variable on the conditioning variables and a treatment indicator.

In my graduate applied econometrics class, we are just finishing the discussion of matching and weighting methods. In my lectures, I make the point that parametric linear regression and matching methods differ in four main ways:

1. Matching relaxes the functional form restrictions inherent in parametric linear regression the way in which it is normally used in applied work, which is to say with each conditioning variable entered linearly and few, if any, higher order terms.

2. Matching focuses attention on the so-called overlap or "common support" condition, which considers whether there are untreated units that "look like" each treated unit in terms of their observed characteristics. With a parametric model, it is easy to rely on the functional form to fill in where the data are absent without knowing that you are doing so. Matching makes that much harder.

3. In one of the "usual" notations, parametric linear regression requires E(U | X, D) = 0 while matching requires E(U | X, D = 1) = E(U | X, D = 0), where U is the "error" term, X are conditioning variables and D is the treatment indicator. This difference in conditions may affect the set of reasonable X. For example, a lagged Y might satisfy the matching condition but not the parametric linear regression condition.

4. As noted by one of the commenters at Gelman's blog, in a heterogeneous effects world, parametric linear regression and matching have different estimands. Matching estimates the impact of treatment on the treated (in the usual case) while parametric linear regression estimates a different weighted average of treatment effects. Angrist has been making this point for a while - see his 1998 Econometrica paper and his new Mostly Harmless Econometrics book with Steve Pischke - but it remains under-appreciated within economics. Perhaps oddly, it is widely understood by sociologists.

A few other points:

1. Thinking about matching as a way of selecting comparison observations is really just a special case of thinking about matching as a weighting estimator. It is a special case because all the weights are integers (or, in the case of single nearest neighbor matching without replacement, they are all one of just two integers: 0 and 1). See equation (10) of Smith and Todd (2005) Journal of Econometrics.

2. One reason to prefer thinking about matching as a version of weighting is that it pushes you away from doing nearest neighbor matching, which the literature pretty clearly shows to have inferior performance relative to its alternatives in terms of mean squared error. For the latest on that literature, see the papers by Busso, DiNardo (get well soon!) and McCrary on McCrary's web page at Berkeley law school.

3. One reason to prefer thinking about matching as an application of non-parametric regression, which is how I teach it in my class, rather than in terms of comparison group selection, is that it makes clear that matching fits much more neatly into our existing stock of econometric and statistical knowledge than it might at first seem.

4. I don't think we fully understand the statistical properties of matching treated as a "pre-processor" in the sense of this paper by Ho, Imai, King and Stuart, which Gelman seems to have in mind in part of his discussion. We do know that doing some statistical procedure on a sample obtained by some sort of matching and not taking note of the pre-processing in the construction of the standard errors will make for misleading inferences.

5. Sometimes you can learn about what conditioning variables are required to make unconfoundedness hold in particular substantive contexts by running experiments. Indeed, to me this is one of the major values of experiments. For this reason, I argue that experiments should often be accompanied by parallel collection of the data required for a non-experimental evaluation designed to shed light on the variables that are, and are not, required for "selection on observed variables" to hold in a given context. For instance, we have learned a great deal about the variables required for selection on observed variables to hold in the context of evaluating job training programs in precisely this way. See, e.g., Heckman, Ichimura, Smith and Todd (1998) Econometrica (gated).

6. Contra Gelman, what you want is not all the variables that determine participation, but rather all the variables that determine both (not either but both) participation and outcomes. A variable that affects participation and not outcomes (other than through an effect on participation) is an instrument. If you have one, you should be using it to do an instrumental variables analysis. You do not want to be in the business of matching on instruments. Also, if you literally had all of the variables that determine participation, you could not do matching, because there would be no common support. Put more prosaically, in such a case, all of the treated observations would have estimated propensity scores of one and all the untreated units would have estimated propensity scores of zero.

7. I really like Gelman's point about the two tribes: those who think unobserved variables are always important, so that selection on observed variables is always wrong enough to lead to substantively important bias, and those who think that selection on observed variables can be true enough in particular, well-motivated contexts to yield reasonable results. I count myself a member of the second tribe, but have many (economist) friends in the first tribe. There is also a third tribe, which I think of as the "benevolent deity" tribe. They believe that whatever variables happen to be in the data set they are using suffice to make "selection on observed variables" hold. This tribe has a lot of members, particularly outside of economics. Indeed, it is probably the largest of the three tribes in the academy as a whole. If you do not believe this, read the chapters in Linda Waite's The Case for Marriage book that survey literatures untouched by economists.

Hat tip: Jess Goldberg

Saturday, October 23, 2010

Canadian census long form (continued)

You can sign a petition regarding either (or both) political interference with Statistics Canada in general or the elimination of the Canadian Census long form in particular. I signed the first of the two petitions.

And you can read articles in Canadian Public Policy on that very topic as well. The pieces on the Census long form controversy are ungated for now. Click on "Show Access Options" and then access them via Metapress.

Tuesday, August 17, 2010

Traffic, NYC, NYT and the mysterious missing denominators

The NYT covers a report on pedestrian accidents on NYC streets from NYC's transportation department.

Here is a bit from the NYT summary:

Pedestrians would be well advised to favor sidewalks to the right of moving traffic — left-hand turns were three times as likely to cause a deadly crash as right-hand turns — and to stay particularly alert at intersections, where three-quarters of the crashes occurred.

Could it possibly be that more accidents happen at intersections because intersections are almost always where the pedestrians are out in the street with the cars?

Some denominators sneak in during a discussion of taxis, but prove a bit complex and confusing:
In Manhattan, about 16 percent of pedestrian crashes that led to death or serious injury involved a taxi or livery cab. Taxis account for only 2 percent of vehicles registered in the city, but at some times of day, they can make up nearly half of Manhattan’s traffic, according to some estimates — challenging the widely held perception of cabbies as the scourges of city streets.
Of course, the correct denominator is vehicle-miles, not vehicles registered, as taxis presumably get driven quite a lot more than the average vehicle, particularly in NYC. Thus, the numbers stated here are consistent with both taxis having safe, professional drivers who pose less of a danger than others and with them being the scourge of the streets. There just is not enough information to know.

Here is the wise transportation commissioner:
“One crash is one crash too many,” said Ms. Sadik-Khan, who said that Monday’s report would help her department “solve the riddle of why people are dying, and where they are dying, in the city.”
Of course, the optimal number of crashes is not in fact zero but rather the positive number at which the marginal cost of avoiding additional accidents equals the marginal cost from doing so. I guess there are no economists and no denominators at NYC's transportation department.

Not a very impressive performance from either the "newspaper of record" or NYC's transportation department.

Hat tip: Jesse Gregory (whose email said (correctly): "seems like the kind of article you love to hate")

Sunday, August 15, 2010

Canadian census long form

The conservative government in Canada is seeking to eliminate the Census "long form" (or, more gently, to make it voluntary). The US did this a few years ago but replaced the long form with the (still mandatory) American Community Survey, so that the issue of having data (almost) uncontaminated by non-random response bias, which to me is the key issue in the debate over the long form, does not arise. The Canadians propose to make the long form voluntary, thus corrupting all the estimates it generates with selection bias.

Here is a fairly reasonable Globe and Mail editorial and here is a somewhat over-heated and remarkably political editorial from Nature.

Here are some tweets from a Canadian government minister who does not, as the saying goes, get it.

Shame on the conservatives in Canada for not being serious on this one. They are an embarrassment to themselves and to Canada and are sending a strong signal that they are not serious about good government. And that matters in Canada, where the analogue to "life, liberty and the pursuit of happiness" is "peace, order and good government."

My pals at Reason also do not distinguish themselves on this one. The fact that the data are not perfect is an argument for improving the data, not abandoning the attempt. More broadly, it seems to me that classical liberals should support govenrment production of public goods, like knowledge, and the evidence-based policy making it supports, particularly given, at least in my experience, that most programs and policies cannot pass a cost-benefit test when rigorous evidence is available.

Regular readers may object that I cannot complain about mandatory jury duty and then not complain about a mandatory census, but the census from mandatory jury duty in three important ways. First, the tax is small. Completing the Census for takes a few minutes. Second, the tax is uniform. Everyone pays the same small tax, rather than jurors who get stuck on long trials paying a gigantic in-kind tax and jurors who escape without being assigned to a trial paying no tax. Third, volunteer jurors are probably a good substitute (perhaps even a superior substitute) for mandatory ones, while voluntary surveys are not a good substitute for mandatory ones, particularly not at the key task of establishing solid benchmark statistics against which all the other voluntary surveys can be judged.

I do think that the Census should aim to minimize the burden of the long form. 20 percent of the population may be much more than is required to get very precise population estimates, even at a provincial or metropolitan area level. In a sense, reducing the population subject to the requirement to complete the long form is what the US did by replacing it by the ACS.

More broadly, given Facebook and Google and all the rest, it seems a bit bizarre for people to worry about the privacy implications of the innocuous questions on the long form.

Hat tip on the tweets to a friend at a government contractor in Canada.

Tuesday, June 29, 2010

Gelman on DiNardo

Andrew Gelman blogs about a piece by my colleague John DiNardo on Bayesian statistics. Not surprisingly, Gelman prefers his own book on Bayesian statistics to John's critique. :)

More seriously, I had not seen this bit by John before and will look forward to reading it (as I look forward to reading Gelman's book at some point). Part of the fun of having John as a colleague is that he thinks really deeply about the philosophical underpinnings of econometrics and statistics.

I do doubt this predictive quality of Gelman's concluding paragraph:

What I suspect--any readers who know DiNardo can ask him directly--is that he is simply unaware of the modern approach to Bayesian data analysis which is based on modeling and active model checking ("severe testing," to use the phrase of Deborah Mayo). I don't expect that seeing my books would make DiNardo a convert to the Bayesian approach, but it might make him realize that practical Bayesians such as myself are not quite as silly as he might imagine.

In my experience, any line of argument that relies on John not having read about something is likely to fail.

Hat tip: Ben Hansen

Monday, May 17, 2010

Statistics 101: a brief rant

So I spent a bunch of time over the past few days reviewing a draft final report on an evaluation that I cannot tell you about because I signed a form promising not to. Maybe I will blog about it when the results are made public and maybe not.

In any case, that is not the point. The point is that the draft had a whole bunch of instances in which an estimated coefficient not being statistically different from zero was equated with the underlying population parameter being equal to zero.

This is wrong!

Large standard errors do not mean that the population parameter equals zero, they mean that the estimate is not very precise. While that is surely disappointing, it is the very truth.

Moreover, even in cases with large standard errors, the preferred estimate - and the maximum likelihood estimate in the case of models estimated using maximum likelihood methods - is the obtained point estimate, not zero.

Classical statistics is odd in a bunch of ways (as any Bayesian will be happy to explain to you at great length) but it is what we have, more or less, so we should get it right.

Wednesday, March 17, 2010

Measuring zombies

Andrew Gelman and a co-author knock the life out of a difficult measurement problem.

Hat tip: Andrew Gelman, of course.

Monday, February 15, 2010

Statistics question

What is the sampling distribution of the r-squared value from a regression?

Saturday, February 6, 2010

The Guardian discovers length-biased sampling

One of the reasons that I cover duration models in my graduate applied econometrics course, even though they are not used that widely in the literature (other than, oddly, in the Netherlands and Denmark) is so that I can cover length-biased sampling.

If you take a random section of spells - hospital stays, welfare receipt, marriage or whatever - you will over-sample long spells relative to their proportion of all spells because they are more likely to be in progress at the time you draw your sample.

The duration literature calls this length-biased sampling and it is another manifestation of the same basic point rediscovered in this column in the Guardian, which explains why your friends will, on average, have more friends than you do, and why, on average, you will be less fit than the other people at your gym (if you have a gym).

Hat tip: Good s**t blog