Showing posts with label evaluation. Show all posts
Showing posts with label evaluation. Show all posts

Friday, July 8, 2011

Why I love the What Works Clearinghouse

From my inbox this morning:
The What Works Clearinghouse (WWC) has released an intervention report this week that reviews the research on Great Books.

Great Books is a program that aims to improve the reading, writing, and critical thinking skills of students in kindergarten through high school. The program is implemented as a core or complementary curriculum and is based on the Shared Inquiry™ method of learning. The program includes both oral and written activities designed to help students think and talk about the multiple meanings of texts. Great Books reading selections are collections of traditional and modern literature. This report focuses on Great Books programs for reading in grade 4 and higher. The WWC identified 36 studies of Great Books for adolescent learners that were published or released between 1989 and 2010. Five studies are within the scope of the Adolescent Literacy review protocol but do not meet WWC evidence standards. Thirty-one studies are outside the scope of the Adolescent Literacy review protocol. No studies of Great Books that fall within the scope of the Adolescent Literacy review protocol meet WWC evidence standards; meaning that, at this time, the WWC is unable to draw any conclusions based on research about the effectiveness or ineffectiveness of Great Books on adolescent learners. Read the full report now at [link].
I like this on several levels. First, and most important, it seriously engages and summarizes the available evidence on a particular educational intervention. The summary is based on precise substantive criteria. Second, it highlights how much time, energy and money is wasted in the education literature on studies that do not meet even very basic standards of evidence. Third, it helps to illustrate how educational interventions can spread despite the lack of any serious empirical demonstration of their effectiveness.

Wednesday, June 8, 2011

Children's stories

Via those portside.org emails that I continue to receive comes Marian Wright Edelman, boss of the Children's Defense Fund, defending Head Start.

My primary concern today is not whether or not Head Start should be cut, it is with how Edelman makes her argument. In particular, Edelman offers two sorts of evidence. The first consists of a moving anecdote about someone called Angelica Salazar:
The colors were brighter than any she had seen before. Shapes, letters, and lots and lots of colors adorned the walls. Around the room, children worked together building high rises with colored blocks and "reading" colorful picture books.

"I had never seen so much color," Angelica Salazar recalls of her first days as a Head Start preschooler in Duarte, California. She remembers her discovery of library books and spending hours curled up on the reading rug. Head Start provided her first formal English instruction. Her parents, who spoke mostly Spanish, enrolled her in the program knowing that their little girl would need to master English to succeed in school.
Anecdotes play a surprisingly large role in policymaking but they are, of course, entirely irrelevant. Anecdotes typically confuse outcomes (Y1 in the usual notation) with impacts (Y1-Y0, or the difference between the outcome with the program and the outcome without it). Put differently, anecdotes ignore the counterfactual of what would have happened in the world without the program. In the case of Angelica, her parents might have found another program, or maybe she would have done just as well with just K-12 education. The difference between outcomes and impacts is remarkably hard for people to grasp. Even in the face of experimental evidence of ineffectiveness, program operators will often insist that they know a program works because of the outcomes they have observed among participants.

The second sort of evidence offered up by Edelman consists of a brief reference to a completely different program:
Early childhood programs significantly increase a child's chances of avoiding the prison pipeline that Angie now studies as a policy expert, and investments in quality early education can produce a rate of return to society significantly higher than returns on most stock market investments or traditional economic development projects.
Edelman does not say what this program might be, but it is presumably Perry Preschool. Jim Heckman and some of his recent students have thoroughly reanalyzed the Perry data, bringing to the discussion a level of rigor and seriousness unfortunately often missing from the advocacy research of the program's developers. When they are done, Perry does indeed pass cost-benefit tests and have a positive rate of return. But Perry is not Head Start. Perry was much more expensive than Head Start, featured different sorts of instructors, served a different and more disadvantaged (particularly cognitively) population and included interventions with parents. There are also important unanswered questions about the extent to which it could be effectively scaled up. Perry is not irrelevant to Head Start, but Edelman is being dishonest by writing as though they are the same thing while ignoring the evaluation record on Head Start itself. On their own, the Perry findings certainly support expenditures on research; any relevance to Head Start funding is speculative.

On the substantive question, the literature on Head Start is mixed at best. Michael Baker of Toronto gave a very nice talk at the Canadian Economic Association meetings last weekend that reviewed the serious literature on early childhood interventions. If you are willing to wait awhile, the talk will be published in article form in the Canadian Journal of Economics. In the meantime, the experimental evaluation of "early Head Start", helpfully summarized at the Institute of Education Sciences' What Works Clearinghouse is a good place to start.

In sum, Edelman's piece has value mainly as an example of an attempt to distort evidence and mislead readers. Her closing appeal to "values", like appeals to "fairness" in similar contexts, is the last refuge of an evidence-avoiding scoundrel. And what of portside.org? Evidently scientific socialism, as they used to call it, has been tossed overboard for the simpler pleasures of sentimental socialism, in which reason and evidence are replaced by heartwarming bedtime stories. Sigh.

Monday, June 6, 2011

State of the Art Lecture from the Canadian Economic Association Meetings

The slides from my "State of the Art" lecture entitled "Putting the Evidence in Evidence-Based Policy" that I gave at the Canadian Economic Association meetings in Ottawa this past Friday are now available on my web page.

I was flattered to be invited and really enjoyed it, though as usual I tried to talk about too much in the time allowed. I was delighted to be introduced by my friend Barbara Glover, who is one of those rare government bureaucrats who is both successful in government and "gets" academics and what they have to offer to the policy process.

One message that I hope got across in my talk is that Canada is way behind the Nordic and Germanic countries in Europe in terms of the data available for policy design and evaluation. This is also true in the US, but the US partially makes up for it by doing randomized social experiments, at least in the areas of labor market policy, education policy and criminal justice policy, while Canada does not.

My talk at the CEAs is a modified version of a talk I gave in Oz two years ago. The associated paper, which covers many of the same issues in textual form (and provides full citations) is on my web page here.

Thursday, April 14, 2011

Evaluating NDLP

In the post by Matt Khan that I blogged about a few weeks ago, he argues that one purpose of a blog is publicize one's own research. I have not done much of that, other than in a very broad sense, but it is probably a good idea.

So, in that spirit, I note this paper, which I recently circulated as an IZA working paper:
The Impact of the UK New Deal for Lone Parents on Benefit Receipt

Peter Dolton
Royal Holloway College, University of London,
London School of Economics and IZA

Jeffrey Smith
University of Michigan,
NBER and IZA

This paper evaluates the UK New Deal for Lone Parents (NDLP) program, which aims to return lone parents to work. Using rich administrative data on benefit receipt histories and a “selection on observed variables” identification strategy, we find that the program modestly reduces benefit receipt among participants. Methodologically, we highlight the importance of flexibly conditioning on benefit histories, as well as taking account of complex sample designs when applying matching methods. We find that survey measures of attitudes add information beyond that contained in the benefit histories and that incorporating the insights of the recent literature on dynamic treatment effects matters even when not formally applying the related methods. Finally, we explain why our results differ substantially from those of the
official evaluation of NDLP, which found very large impacts on benefit exits.
This is an old paper. It got started in the late 1990s as a result of the two of us being on the technical advisory panel for the official UK government evaluation of the NDLP. The results of that evaluation were so positive that even the government agency sponsoring the evaluation did not quite believe them (a high bar indeed!) and so Peter and I, along with his gradual student Joao Pedro Azevedo who is now at the Wild Bank, were retained to reanalyze the data.

Filled with hubris, we expected that if we just redid the analysis and tweaked the methods a bit, the results would change dramatically. This turned out to be wrong. Neither varying the matching method nor worrying a lot about survey non-response moved the estimates very much at all. Changing the outcome variable from "leaving income support within six months" to a monthly measure of benefit receipt did matter some. We document these findings in this report for the UK Department of Work and Pensions.

We continued to pursue the mystery of the oddly high impacts. The paper reflects the additional work we did after writing the report for the DWP. In the end, using just the administrative data, we are able to get the impact estimates down to levels that are likely still a bit too high, but at least within the range of plausibility suggested by the literature. Our analysis reinforces the message from some of my earlier work with Heckman and others regarding the importance of flexibly conditioning on pre-program outcomes. We also indirectly show the value of recent developments in the literature on dynamic treatment assignment, as in the important paper by Barbara Sianesi (2004) Review of Economics and Statistics. We were surprised to find that questions about attitudes toward work do surprisingly well as conditioning variables, while variables related to local labor markets matter very little, contrary to the findings using the JTPA data in Heckman, Ichimura, Smith and Todd (1998) Econometrica. This latter point merits further investigation: when do local labor markets matter in evaluating active labor market programs and when do they not?

Given the heavy UK content, we sent the paper to the Economic Journal, which is the journal of the Royal Economic Society.

Regular readers will note that our estimator in this paper is drawn from the curio cabinet. That is to say, we rely on nearest neighbor matching. The reason for this is essentially path dependence. When we started the work in the early noughties, the gain in speed from using nearest neighbor matching in preference to kernel matching - inverse propensity weighting was not on the radar screen at that point - was large enough that it seemed the correct choice. Once we started down that path, we never quite got around to changing estimators to something better. We'll see what the referees think about that.

Doctoral fellowship at MDRC

This would be a lot of fun!

MDRC is one of the very best social science research organizations, with a primary focus on random assignment evaluations of social and educational programs.

4/15/11: Link fixed.

Sunday, March 27, 2011

Laundering studies

A nice piece from EduWonk on how studies get "laundered" when some credulous / differentially not competent person at NYT or WaPo is unable to sort out, and ignore (or even, perish the thought, call out) bad studies, which then become reputable because they have been discussed in a major newspaper.

The discussion of achieving "mixed evidence" by vote counting in which weak studies receive the same weight as strong ones is also useful.

EduWonk presents the issues in terms of the literature and policy discussion surrounding Teach for America, but the same issues apply in nearly every area of policy-related empirical work. They result from a combination of advocacy groups producing output that looks like scholarly research but is not with newspaper reporters, columnists, bloggers and others who lack the quantitative skills and context knowledge required to correctly weigh the evidence.

Saturday, March 5, 2011

The non-market time of the poor and cost-benefit analysis

I just learned about a new paper yesterday that makes a point that I think is quite important. Here is the abstract:
Benefit–cost analysis is used extensively in the evaluation of social programs. Often, the success or failure of these programs is judged on the basis of whether the calculated net benefits to society are positive or negative. Almost all existing benefit–cost studies of social programs count entire increases in income accruing to participants in a social program as net benefits to society. However, economic theory implies that the conceptually appropriate measure of the impact of a government program on any group of individuals is the net change in their surplus (or economic rent), rather than the net change in their income. For example, if a social program causes increases in income by increasing work hours, then the lost nonmarket time that accompanies these increases has value that needs to be counted as a cost when assessing the merits of that program. In this paper, we develop a methodology for incorporating lost nonmarket time into benefit–cost analyses of social programs. We apply our methodology to the Self-Sufficiency Project (SSP), an experimental welfare-to-work program tested on a pilot basis in two provinces in Canada during the 1990s. We find that if losses in nonmarket time are ignored, SSP yields a substantial positive net benefit to society. However, if losses in nonmarket time are taken into account, the net societal benefits are greatly reduced, even becoming negative in certain instances. We conclude that future benefit–cost analyses of social programs must take effects on nonmarket time into account in order to give a more accurate picture of the net benefits of the program.
The full citation is:

Greenberg, David and Philip Robins. 2008. Incorporating Non-market Time into Benefit-Cost Analyses of Social Programs: An Application to the Self-Sufficiency Project. Journal of Public Economics 92(3-4): 766-794.

You can find a gated version here.

Valuing the non-market time of the poor means taking the economics seriously in doing the cost-benefit analysis, but it also means that fewer programs will pass cost-benefit tests.

Wednesday, January 12, 2011

Data, evaluation and the quality of public policy

This new working paper describes an evaluation data set that researchers can use to study the effects of active labor market policies in Germany. The data set combines administrative data drawn from multiple sources with innovative survey data.

There is nothing like this available in the US or Canada. Under the guise of privacy concerns, the political process in the US and Canada avoids assembling the sorts of data that would allow very high quality non-experimental evaluations (as well as descriptive analysis useful for understanding program operation and informing program implementation and design) of active labor market programs. Part of this is, at least in the US, due to the fact that both political parties have strong, but different, prior beliefs, about the effectiveness of such programs. To many democrats they are obviously effective (how could more schooling be bad?) while for republicans they are obviously ineffective (how could bureaucrats increase anyone's employment chances?) thus the demand for quality empirical analysis is low regardless of who is in power.

What we might call the "data gap" (a play on the historical "missile gap") has the effect of leading US researchers, at the margin, to spend their time working on non-US data, as with Dale Mortensen's ongoing research project in Denmark and Sandy Black's ongoing use of the Norwegian register data. Now, to be sure, no one actually moves outside the US because the salaries are much too low, but they do change their travel plans and their research agendas. In some ways this is good, because it leads to more interactions between North American and European researchers. However, the US is a big country, and policy unguided by serious evaluation in the US affects a large number of people and one of the world's most important economies.

There are real, low-cost opportunities for dong some good here. Will anyone in DC pick up the ball and run with it?

Wednesday, September 22, 2010

Teacher performance pay

Here is the WaPo piece on the teacher performance pay experiment that I described in an earlier post and that is being presented at the Ford School this afternoon.

The experimental design here is solid but there are some issues of interpretation that stem directly from the nature of the study design.

1. Only volunteer teachers were randomized. Whether one should expect the effects of performance incentives to be higher or lower for volunteers than non-volunteers is not clear a priori but using volunteers makes the politics of doing a random assignment study much easier. The importance of this issue also depends on the take-up rate, which as I recall was higher than you might think but not 100 percent by any means.

2. This is a temporary program. That means that incentives to invest in being a better teacher are limited, as they pay off only in the short term. The numbers at play in the study are not trivial but also not large relative to, say, the cost of going back to college to get a subject area degree.

3. As Rick Hanushek notes in the article, by design the study looks only at current teachers. One important effect of an ongoing regime of performance pay might be to change the mix of people who become teachers in good ways. This study cannot pick up that effect.

4. The treatment is just performance pay. There is no mentoring or other treatment. If teachers are going to figure out how to do better, they are on their own with the literature. Knowing something about the literature, I know that is problematic. But it is not clear that combining performance pay with the sort of additional treatments mentioned in the article, such as mentoring, would work. For example, the randomized trial that the Dept. of Education funded on additional mentoring for new teachers did not produce clear positive results.

I should note in relation to the last point that things are very different in many developing countries. There the effort margin is often quite important. In some countries, for example, teacher absence rates are shockingly high. Performance pay might well get them to show up, and showing up would likely improve test scores.

A side point: why is Rick Hanushek the only person quoted who is identified as having political beliefs? The Hooover Institution where he works is labeled "conservative leaning" while the fellow from the teacher labor cartel is not given any politics, nor are the people from the federal government, nor is anyone else.

Hat tip: Dann Millimet

Friday, August 6, 2010

Idea for new Institute for Education Sciences randomized trial

They should evaluate this book:
Hot X: Algebra Exposed

Description: New York Times bestselling author Danica McKellar tackles the toughest math class yet: Algebra! In her two bestselling books, Math Doesn't Suck and Kiss My Math, actress and math genius Danica McKellar shattered the "math nerd" stereotype by showing girls how to ace middle school math-and actually feel cool while doing it! Sizzling with Danica's trademark sass and style, Hot X: Algebra Exposed tackles algebra: the most feared of all math classes and the most common roadblock to high school graduation. McKellar instantly puts her readers at ease, showing teenage girls-and anyone taking algebra-how to feel confident, get in the driver's seat, and master topics like square roots, polynomials, quadratic equations, word problems and more . . . without breaking a sweat (or a nail). Danica provides illuminating, step-by-step math lessons combined with reader favorites like personality quizzes, popular doodles, real-life testimonials, and stories from her own life, so girls feel like she's sitting right next to them. As hundreds of thousands of girls already know, Danica's irreverent, light-hearted approach opens the door to higher grades and higher test scores. Now, with Hot X: Algebra Exposed, the scary veil of algebra is finally lifted, making it understandable, relevant and maybe even a little (gasp!) fun for girls.
I am eagerly awaiting the Jonas Brothers' new book on chemistry.

Hat tip: Charlie Brown

Monday, April 26, 2010

Evidence based policy in Oz

The volume containing the paper I wrote based on my remarks at the conference on Evidence Based Policy organized by Australia's Productivity Commission is now available on-line. My piece, which is an update of an older paper that my friend Arthur Sweetman and I prepared for an HRSDC conference many years ago, is Chapter 4, and is entitled "Putting the Evidence in Evidence-Based Policy".

The conference itself was great fun, and I was impressed with the general caliber of both the political and academic participants.

Thursday, March 4, 2010

Experimental evaluation of Head Start

The official reports from a random assignment evaluation of Head Start were released in January; go here for the Executive Summary. Head Start is the federal pre-kindergarten program for disadvantaged children in the US. It dates back to the War on Poverty of the 1960s. It is one version of what you get when you try and scale up, but not spend too much money on, high intensity interventions like the famous Perry Preschool project in Ypsilanti that receives so much attention in the literature and whose long term experimental impact estimates are often held up as justification for expansion of programs such as Head Start.

I have not had a chance to read over this carefully but the results are pretty important. The study has three key design elements: (1) random assignment of access to Head Start; (2) the study sites are a random sample of Head Start sites around the country, which suggests strong external validity; and (3) a large amount of substitution into other types of center-based childcare and also a surprisingly large amount of control group cross-over into Head Start. There is also a non-trivial amount of treatment group dropout.

Factor (3) is important in interpreting the rather lackluster findings. The study is comparing the offer of Head Start access to the best alternative (which in some cases turns out to be a different Head Start center not in the experiment). The study is consistent with Head Start having a positive effect relative to the child staying home with the parents, but not much in the way of effects relative to the alternatives chosen by the parents in the absence of access to Head Start.

This study should slow the rush to spend large sums on these programs and hopefully will encourage a systematic program of research aimed at learning what types of early education programs have persistent impacts that cover their costs for this age group.

Wednesday, September 16, 2009

The ethics of random assignment

Tyler Cowen at MR asks for information related to the ethics of random assignment.

I am going to put my comments here as not everyone who reads this blog may delve into the comments (or even, gasp, the posts) at MR.

Most of my experience with these issues is in the context of randomized evaluations of social and educational programs rather than the clinical trials that are likely to be of greater interest at NIH.

I have the following comments:

1. Ethical concerns about random assignment in the social policy world are nearly always fake. They are a nice way of saying that the person has some interest in there not being compelling evidence on the effectiveness of the treatment being evaluated.

2. It is easier to sell random assignment for demonstration programs, which by design are going to leave many people who want service without it in any event, than on-going programs. With on-going programs it is easier to sell random assignment for programs that are capacity constrained, so that services will have to be rationed in some way even in the absence of random assignment. Random assignment then just becomes a way of doing the rationing that happens to yeild an informative byproduct. This is the line that was popularized by Judy Gueron during her days running MDRC, which did many of the welfare-to-work experiments in the US in the 1980s and 1990s.

3. An argument that is not often made, but which in my experience is quite convincing even to non-classical-liberal types is that there is a competing ethical claim associated with using tax money to fund the provision of services without solid evidence of their efficacy. Just as the shareholders of a firm would rightly be upset if it undertook major investments without doing its due diligence beforehand, taxpayers should rightfully be upset when government spends money on programs without doing the (in many cases) simple, obvious and relatively inexpensive things required to determine their efficacy, or lack thereof.

Saturday, June 20, 2009

NYT discovers selection bias

A not-too-bad report from the NYT on the potential for selection bias in studies that show health benefits from alcohol consumption.

My comments center on this bit:

Meanwhile, two central questions remain unresolved: whether abstainers and moderate drinkers are fundamentally different and, if so, whether it is those differences that make them live longer, rather than their alcohol consumption.

Dr. Naimi of the C.D.C., who did a study looking at the characteristics of moderate drinkers and abstainers, says the two groups are so different that they simply cannot be compared. Moderate drinkers are healthier, wealthier and more educated, and they get better health care, even though they are more likely to smoke. They are even more likely to have all of their teeth, a marker of well-being.
What matters is not whether the mean characteristics differ but whether there is what in the technical literature is called "common support", which just means overlap. Are there some non-drinkers who have the characteristics of moderate drinkers? If so, that is enough.

More important than overlap is, of course, measuring all the relevant confounders. A different, and perhaps better, way to frame the article given the likely impossibility of doing a major random assignment study on moderate alcohol use would have centered on the importance of having a large data set that contains information on all the variables that might confound the effect of alcohol on health.

The NYT piece, unfortunately, does not really make it clear how good a job the existing literature does at including all the possible confounders. I suspect the answer is "not very well" and that there is room for improvement here even without a randomized trial, though obviously that would be helpful if one could be fielded.

Hat tip: Jessica Goldberg

Saturday, April 11, 2009

Health campaigns that failed

Cracked magazine summarizes the evidence on some failed public health campaigns; sadly, their summary is much more honest that you might get from some quarters or the government or the public health establishment.

Monday, February 2, 2009

IES TWGs in DC

I was in DC last week for two days to attend meetings of two Technical Working Groups for evaluations being funded by the Department of Education's Institute for Education Sciences (IES). These techincal working groups include staff from the evaluation contractor and IES as well as outside experts on methods (that is my usual role) and on the subject area. The outside experiments are usually mostly academics, though they are sometimes program operators or, in the case of IES evaluations, school district or teacher union officials.

On Wednesday, I went to the TWG for the evaluation of mandatory random drug testing being done by RMC Research in cooperation with Mathematica. This evaluation is at the stage of having preliminary results (which I am sworn to secrecy about) so we discussed various statistical issues related to the analysis as well as issues of presentation and focus and some secondary analyses that it would be worthwhile to undertake. The IES page on this evaluation is here.

A very interesting issue here, which we discussed at some length, is whether the key dependent variable is any drug use or frequency of drug use. This issue might seem minor, but in fact it hinges importantly on what one sees as the point of the drug testing program. If the point is to get students down to zero use, then a dummy variable for zero use is the appropriate dependent variable to highlight. In contrast, if the point is to discourage frequent use, say by moving daily or almost daily users down to occasional weekend users, then a categorical variable measuring frequency of use becomes the primary object of interest. A dummy variable for zero use versus non-zero use completely misses changes in intensity that do not lead to abstinence.

This panel was great fun, in part because I was the only economist among the experts in attendance (some of the IES folks and the consultants are also economists). We established at the last meeting that I was the only person in the room who knew what "420" meant, which I thought was kind of amusing. In any case, the other experts on the TWG are a particularly bright and outspoken lot of psychologists and such, so the discussion was great fun and very stimulating.

On Thursday I attended the TWG for the evaluation of a treatment designed to move high performing (as measured by value added in test scores) teachers to low-performing (as measured by test score levels) schools. This evaluation is earlier in the game, as the research team is just completing a pilot study of the program in a single district, and just beginning the broader evaluation in multiple districts. Though the basic design is largely set, there was lots to talk about here as well, including how likely one thinks spillovers are from high performing teachers and how long it might take any such spillovers to show up in the data on test scores in other classrooms. This has implications for what you measure and how you measure it and also for broader issues of power.

Thursday, January 22, 2009

Counterfactuals, musically



Hat tip: Jessica Goldberg

Sunday, January 11, 2009

Mom, I want to be an evaluator

I like this post from Chris Blattman a lot!

Here are some thoughts:

First, it provides good advice on what human capital you should accumulate if you want to be an evaluator. I would just add a bit more refinement to Chris' recommendations. If you want to do evaluation in development, find a place that is strong in micro-development and in labor. Princeton and Michigan are examples here. If you want to do domestic educational evaluation, find a place that is good in labor and/or education and, ideally, where the economists talk to the ed school people and the ed school people are worth talking to. Harvard and Michigan come to mind here. I would second his recommendation regarding obtaining some work experience in the real world of evaluation but would broaden it to include working at an evaluation consulting firm like Mathematica, Abt or RTI.

Second, it really is true that getting a reliable impact analysis is just the first step. There is always interest in improving even the most effective programs (where effective programs are a rather modest subset of all existing programs) by tweaking the design. There are also important issues of portability. For example, just because Progresa "works" in Mexico does not mean that a Nicarguan or Brazilian version will obtain the same results. Knowing how and why a program produces the impacts it produces makes addressing these questions of program modification and program replication in other contexts a much surer business. Of course, as a byproduct learning about how and why programs work adds to our general stock of social science knowledge in a way that isolated black-box impacts estimates do not. This point is emphasized in, e.g. the Heckman and Smith (1995) piece in the Journal of Economic Perspectives.

Third, I would note that economists sometimes forget that there is more to evaluation itself than obtaining credible impact estimates, difficult though that often is even within the context of an RCT. If you pick up one of the standard non-economics texts on evaluation, such as the nice book by Shadish, Cook and Campbell, you will find that surprisingly little space (probably too little, but that is a different post) is spent on impact evaluation. What you find are are sections on topics such as implementation fidelity (does the program as implemented match the program as designed) and process evaluation (how does the program operate in practice) that economists tend not to talk too much about.

Finally, cost-benefit analysis, which in economics traditionally falls under the heading of public finance, is also an important part of a serious evaluation. Very few evaluations do a very good job with this; a useful exception in the labor economics world is Mathematica's experimental evaluation of the Job Corps program. Even if your career focuses on impact evaluation, having a solid foundation in the other facets of evaluation will allow you to do your job better. You will be able to interact more effectively with people from other backgrounds and you will find that, for example, the output produced under the heading of a process analysis is very valuable in the course of an impact analysis. An excellent examples of a (relatively quantitative) process analyses in the labor economics context is the Kemple, Doolittle and Wallace (1993) report produced as part of the National Job Training Partnership Act (JTPA) study.

Hat tip: marginal revolution

Tuesday, July 29, 2008

Multiple comparisons

Economists do a pretty bad job, in general, of being careful when making what the literature calls multiple comparisons. That term encompasses situations where, for example, the researcher has quite a lot of dependent variables and is considering impact estimates for all of them. Of course, if there are 100 independent outcomes, then we would expect statistically significant impacts on five (or ten depending on your alpha) of them even in the world of the null where the treatment has no effect on any of them.

The social science and statistics literatures contain a number of ways to address this issue somewhat formally, including various procedures for adjusting significance levels, such as the Bonferroni adjustment, and dimension reduction techniques applied to the outcome variables.

This recent report by Peter Scochet of Mathematica for the Institute for Education Sciences provides a useful introduction to the (surprisingly tendentious) literature.

Saturday, July 19, 2008

Evaluating Amtrak

An example of the wise and practical American approach to nationalized industries can be found here.

The video is hilarious but may not be work-appropriate at particularly fastidious workplaces.