Friday, April 09, 2010
More on Authorship
1) How does the community properly assign credit for 60-40 papers? Should we use author ordering or some other mechanism to assign credit?
2) What about advisors who do minimal to zero work but put their name on the paper?
3) At what point has a person involved in the project done so little work that they should not be included as an author (by either withdrawing willingly, or possibly by being told "you're not an author"). (I think of this as separate from the "advisor" issue.)
Let me start with item 1, assigning credit. I promoted the approach used in theory (derived, apparently, from mathematics) of alphabetical order, claiming credit comes out through things like letters and who gives the talk, and is determined more clearly over the course of a career. Many question this; indeed, many other fields use entirely different systems. Many fields use author order to signal the level of contribution in some way, so that being "first author" has significant meaning. At the extreme, the journal Nature, for example, suggests that author contributions should be fully specified in each article in their guide to authors:
"Author Contributions: authors are required to include a statement to specify the contributions of each co-author. The statement can be up to several sentences long, describing the tasks of individual authors referred to by their initials."
Graduate student and postdocs, in particular, are more concerned with systems that clarify credit, and this is understandable. They have short career track records, and want a job; making sure that they get their proper credit often seems, to them, quite imperative.
I'd like to defend the alphabetical, no-explicit-credit-assigned system, and then provide a couple of stories. (If you find that indulgent, you can skip the stories.)
One philosophical approach is to try to start from a blank slate. Forget about your current situation, and how your field does things. Your starting point is that you're just starting a career in science. What sort of system do you want to use? I'd argue you'd want to use a system that would lead to long-lasting, productive collaborations; that would have minimal overhead; and that would still provide meaningful ways of calibrating people over appropriate time periods. I think pure alphabetical does that. It removes the need to fight over (or even discuss) who contributed exactly what, leading more easily to frequent and repeated collaboration. To be clear, I have a strong bias: collaborations, I think, are great for scientific production, and on the whole make research much more fun. Alphabetical order is clearly easy. And while it's weak on allowing someone to find out how much each individual author contributed to a specific multi-author paper, over the course of several papers, I think the calibration works, especially when augmented with additional information such as letters in job searches and promotion cases. Further, it's not clear that other systems are really stronger in terms of assigning credit. Authors can disagree on contributions -- how does this get settled, and what does it do to future collaborations; in multi-author situations where order ostensibly matters many advisors will game the system, for example by putting students first regardless of their contribution in order to prep them for the job market or out of professional courtesy; and it's not clear how, for example, to value different types of contributions, like ideas vs. data collection and analysis. My bias is that the blank slate scientist starting their career would pick the alphabetical order system.
I have at least one data point for this conclusion: myself. (Here's where the stories start.) In graduate school, a bunch of us students got together and wrote a paper. This was a case where I was definitely the 60 author, and I thought it would be best if I was first author. The other students didn't object, but since I knew it wasn't standard for theory, I asked my advisor. (He wasn't a co-author for this paper, so his view was not biased in that regard.) He told me it was my choice, but that I needed to recognize the following: I would possibly get more credit for this paper, but, from then on, I would have adopted a system where, for every paper, I'd have to face the possibility of constructing the author order with my co-authors. Did I want to have that discussion for every paper down the line? I went with alphabetical order and have never looked back. I always recommend alphabetical order, although when I work with people in other areas I do defer to whatever system they want to use, and tell them they can put me wherever they like in the ordering. (It is true that, with tenure, one can care much, much less about such things.)
On the other side, another story. When I applied for my CAREER grant, apparently I was on the borderline, and it took quite some time to get the final word. I asked the NSF officer for feedback -- especially in case I needed to resubmit. (Apparently, enough money came through in the end to fund me.) One thing he said was that a lot of my work had been co-authored with very talented people, and it wasn't clear what my contributions were. This was a case where, obviously, there were no recommendation letters to draw from. Still, I was offended then by the comment, and looking back I still find it ridiculous. At that point, I'd written multiple papers with these other authors (who were not my advisor) -- clearly they thought I was contributing something worthwhile. And why was the assumption that they were the 60 contributor, instead of me? It's not clear that using author ordering would have helped in this case, or that such cases are at all frequent. But it does help me understand alternative points of view on the underlying question.
Wednesday, April 07, 2010
60-40 papers
60-40 papers aren't at all abnormal, and I've done enough papers not to let it bother me. When I'm the "40" author, I usually try whenever possible to do what I can to help even things out, for example in the writing/editing/revising stages; when I'm the "60" author, I recognize that the other authors have contributed, and the paper wouldn't be what it is without them. I've had amusing discussions with one co-author where we ended up admitting we both thought we were the "40" author for the paper we were writing. That was a collaboration that lasted for several papers; apparently, we both thought we were getting a good deal. I don't think I've been in many collaborations where multiple authors thought they were the "60", but my guess is those could be problematic.
Fan Chung has a nice page up with advice for graduate students that I think puts the 60-40 issue in perspective. At the end, under research collaboration:
What about the division of credit?
-- In math, we use the Hardy-Littlewood rule. That is, authors are alphabetically ordered and everyone gets an equal share of credit.
-- The one who has worked the most has learned the most and is therefore in the best position to write more papers on the topic.
-- If you have any bad feeling about sharing the work or the credit, don't collaborate. In mathematics, it is quite okay to do your research independently. (Unlike other areas, you are not obliged to include the person who fund your research.) If the collaboration already has started, the Hardy-Littlewood rule says that it stays a joint work even if the contribution is not of the same proportion. You have a choice of not to collaborate the next time. (If you have many ideas, one paper doesn't matter. If you don't have many ideas, then it really doesn't matter.) You might miss the opportunity for collaboration which can enhance your research and enrich your life. Such opportunity is actually not so easy to cultivate but worth all the efforts involved.
I'd just add a bit to this. Usually the "60" author will, actually, get more credit in various ways: usually they're the one to give the talk on the paper, for example. (It can also come out in letters when really needed.) And it's not so clear that a string of 60-40 collaborations with one author repeatedly being the 60 is so bad; without the 40, the research or the paper might not ever get done! Good collaborations are indeed enriching. To some, particularly graduate students, this approach and attitude might seem strange, but I recommend considering Fan's suggested understanding of collaboration.
To all the co-authors out there who have been the 60 to my 40, I appreciate your putting up with me. And to all the co-authors who have been the 40 to my 60, as long as we had a good time working on the paper, no worries, and thanks!
Monday, April 05, 2010
Sexual Harassment Policies (Yale v. Harvard)
I was all ready to start looking down my nose at the competition for being slow to adopt what are in my mind obvious rules to have, but decided to check Harvard's policy first. (Always a good idea.) Harvard's policy, arguably, isn't even as strong as Yale's old policy. (Harry Lewis will, I imagine, correct me if I am mistaken in my interpretations or usage of documents.) The relevant information seems to be here. The policy description includes the following, under the heading UNPROFESSIONAL CONDUCT IN RELATIONSHIPS BETWEEN INDIVIDUALS OF DIFFERENT UNIVERSITY STATUS:
"Officers and other members of the teaching staff should be aware that any romantic involvement with their students makes them liable for formal action against them."
This seems to suggest that faculty can't have "romantic involvement" with their students, but some old letter to the Crimson suggests that the wording is much weaker than that (the article is here, the letter is here). Strictly speaking (according to the letter), the wording seems to suggest that faculty members involved with students face the risk of a the student filing a sexual harassment/unprofessional conduct complaint; but if the relationship is brought to light by a third party, there's no (apparent) cause for disciplinary action. IANAL, but this seems like a possible interpretation; I'm not sure what the current interpretation is here at Harvard.
Indeed, later on the policy states:
"Amorous relationships between members of the Faculty and students that occur outside the instructional context can also lead to difficulties."
The rest of the paragraph suggests potential problems if Faculty engage in "romantic involvement" with students who they are not directly teaching, but seems to make clear (by my reading) it's not forbidden in any sense.
I've certainly heard arguments in the past that such rules shouldn't exist. I can even see that there are potentially complicated lines -- should a professor in the Faculty of Arts and Sciences not be allowed to date a Harvard Law student? (Extra credit: why or why not?) But given the potential for abuse (both intentional and unintentional) of the power relationship, I'm unapologetically on the "no faculty - undergraduate romance" side. Or, as it says in the Yale Alumni Magazine:
'An imbalance of power forms the rationale for treating Yale College students differently from their older counterparts. Undergrads, the revised handbook says, “are particularly vulnerable to the unequal institutional power inherent in the teacher-student relationship and the potential for coercion, because of their age and relative lack of maturity.” '
Duh. Good for Yale.
NSF Review Issues
Also, this year, I have been asked (more than once) to review a single proposal "off-panel" (that is, I didn't serve on the panel that the proposal was part of). I can't recall having been asked to do this before, and wonder if there's a policy change behind it or if it's business-as-usual and I'm only now noticing it. I certainly don't mind -- I'm more than happy to help the NSF, and even more happy if I can do so without having to travel to DC. On the other hand, I worry that this approach might cause the same sort of problems that can occur with subreviewers, such as consistency across reviews.
Friday, April 02, 2010
Energy Sustainability
The talk was based on the book. David's starting point is the question, "What would we have to do to move to a world where we weren't using fossil fuels?" (The "we" he's talking about is usually the UK, but it applies elsewhere as well.) He then takes a truly scientific approach. He considers various possible renewable energy sources (wind, solar, biomass, tides), and estimates things like their energy output per unit area. Based on these calculations, he figures out how much land would be required. So, for instance, if you were willing to cover 1/2 of Britain with windmills, things might look OK, but that's not a very likely possibility. He also considers the demand side of the equation, and what might feasibly be done there.
The book (and the talk) are not overtly political. Whether you believe in global warming or not, the question of sustainable energy is important -- for national security concerns, you might not want to be dependent on getting your energy from, for example, oil-rich countries. His book is not about the politics; rather, he tackles these questions as a scientist, producing the numbers that are needed for intelligent, reasoned discussion and debate on the issues. That might sound dry and, possibly, boring, but not in David's hands. He's blessed with a fine wit and a charming style that comes out in the book and even more so when he's speaking. (His slide showing a collage of posters from places protesting the introduction of windmills into their community, for instance, received a lot of laughs.)
David has done a truly rare thing as a scientist, writing a book firmly about the science of an issue of current import that people are actually reading and that is raising the level of debate. I admire his courage in taking on a challenging assignment, and hope his work helps lead to the positive changes he is looking for.
Tuesday, March 30, 2010
Two Inspiring Posts
Inspiring me to blog, that is....
A third inspiring post by way of Muthu. Tim Roughgarden has won the Grace Murray Hopper Award, and Bellare and Rogaway have won the Paris Kanellakis Theory and Practice Award. Congratulations all around!
Wednesday, March 24, 2010
Reading a Research Paper
There are other pages and places with more advice on the topic, so you might want to look around. Searching on phrases like How to Read a Research Paper (and other variants, like How to Read a Technical Paper) will yield plenty of pages. I like the "What to Read" section at the bottom of this page from Jason Eisner -- creative Web search and forward/backtracking references are, I think, often underappreciated skills.
Thursday, March 18, 2010
Google-Viacom Documents
Tuesday, March 16, 2010
More on SIGCOMM
Matt Welsh, perhaps at least partially inspired by being on the SIGCOMM PC as well, has suggested an approach for dealing with the large quantity of awful submissions by charging authors to submit; Suresh disagrees.
Interestingly, the scores (and reviews) on my papers were generally consistent across the line, with a rare exception or two. Since there's always some overenthusiastic anonymous commenter who thinks its important to call SIGCOMM an insider's club whenever I blog about it, I'll repeat that I'm not aware of belonging to any such club, and once again, the reviews I see from others not only make sense, they match my own opinions, which I view as independent, to a striking degree.
Monday, March 15, 2010
Kindle Textbooks
I suppose I'll see in my next book sales statement if any Kindle copies were sold. In general, the Kindle version seem to go for just a few dollars less; Algorithm Design is a big exception, with the Kindle edition going for over $25. less than the hardback counterpart. Sometimes, it looks like you can buy new copies of the book less than the Kindle price (usually from Amazon third party dealers). I don't own a Kindle (yet), and I wonder how they would be for textbooks. The textbooks I've listed above I'm happy to have on my shelves for reference. I suppose if I had them in a universal, pdf-like format that I could access and make use of essentially anywhere, I'd be happy with that too. Particularly if I had capabilities like search available. But I wouldn't want my copies of the book tied to a particular piece of hardware. If I lose my Kindle, do I lose my books? That's fine for disposable books -- and probably for many students many textbooks fall into that category. (Use them for a semester, then forget about them.) It wouldn't be fine for me for these books.
Can anyone comment on the Kindle textbook experience? I'm interested generally, and in the particular issue of the "permanence" of books that might be references one wants to keep. Someday soon, I'll be getting one of these (or another e-reader), and it would be useful to know whether it's currently worth moving to a system where I try to keep important texts in an electronic, rather than paper, format.
Saturday, March 13, 2010
And More Fun News....
The second is much more amusing. Apparently, the Harvard Robobees project has made #1 (that's right, we're number 1!) on Sean Hannity's list of the 102 worst ways the government is spending your tax dollars. Now, I try to stay apolitical on this blog, but I have to say, I'm impressed by Sean Hannity's lack (well, actually, more like a complete absence) of acumen in understanding the nature of scientific research. A look at the Robobees home page, I would think, would certainly suggest that there's important scientific and engineering questions underlying the long-term challenge of building a robotic bee. Of course, maybe the Hannity camp just objects to the government spending money on science generally, I don't know. I'll go on record as suggesting that nobody from the Hannity camp bothered to look at the Robobee home page.
Good News for Computer Science Majors
Thanks to Stuart Shieber for the link.
Wednesday, March 10, 2010
A Conflict Question
Suppose you're the PC chair, and someone who has submitted a paper asks you NOT to have the paper reviewed by a specific person on the PC. Do you honor that request?It's an interesting question -- that I hope others will comment on -- though as a default my answer would be yes. I certainly have had run-ins of sufficient severity with various people through the years that I would not want (and would likely ask) for them not to review my papers if the issue came up. Looking at it from the other end, if I am on a PC and those people submit a paper, I make sure not to review them. (Usually it is sufficient simply to rank them low on my list of desired papers, but I have also told PC chairs in advance I would not review certain papers if they seemed likely to head my way.) It is not that I actually think I couldn't give a fair review; it's that I think it's inappropriate, in such a situation, for me to give a review in the first place. If as a PC member I have the right (actually, I would say, a responsibility) to refuse to review a paper under such circumstances, it seems fair that a submitter can ask for a specific PC member to not review a paper as well.
Context does matter, though. In the networking conferences I have served on, this is standard -- PC members and submitters are expected to list their conflicts. Indeed, one issue that seems to have arisen lately is that there is suspicion that some people submitting papers are abusing this right, listing people as conflicts when they are not because they are known to be "challenging" reviewers. While I'm skeptical this sort of gamesmanship gains anything (challenging reviewers are usually calibrated appropriately at the PC meeting), it is a concern that once you open the door to such requests, you may need to make sure the privilege isn't abused.
For theory conferences, where many people seem painfully unclear on what "conflict of interest" even means, I'd grant such a request as a matter of course.
Tuesday, March 09, 2010
Congratulations to Chuck Thacker
I had the great pleasure of getting to know Chuck while I worked at DEC SRC. He's a character, a tinkerer, and a great and curious mind. I think recognizing his work -- the Alto -- is a great choice.
EC Papers Up
I was surprised in the acceptance letter to find that there were 45 acceptances out of 136 papers -- an acceptance rate of about 1/3. (Compare with WSDM.) This makes it one of the less "competitive" CS conferences I know of, although a little research shows this is a bit unusual -- last year the numbers were 40/160 or so, so they accepted more papers and had fewer submissions this year. Is that a trend in the making or an accident of timing this year? Also, while I'm an EC newbie, the list of papers looks very interesting, with plenty of top-tier names. "Competitive" or not, I'm expecting high quality. I'm really looking forward to it -- and not just because its location makes it remarkably convenient for those of us in the greater Boston area.
Sunday, March 07, 2010
Carousel (NSDI Paper)
Here's the main idea. In IPS systems, the logger can get overwhelmed during an attack. Bad sources (with other related info) need to be recorded for later examination, but there's only a small amount of on chip memory to buffer bad sources, and only a small about of bandwidth (compared to the rate data is coming into the system) from the memory to the more robust recording infrastructure. How do you get all, or almost all, of the sources?
In our model, bad sources are hitting the system repeatedly -- the denial of service attack setting. On the plus side, this means you don't have to log a bad source the first time it appears - the assumption is it will come back again. On the negative side, you want to avoid sending the same source to the recorder multiple times, as it wastes your small bandwidth. (Dups can be removed at the recording side, though.)
The baseline solution is to just grab a bad source to record for the buffer as soon as you have room after you send one out. We consider a random model -- there's a bunch of bad sources, and the next one to appear is random from that set -- that shows that this approach is bad; by connecting it to the coupon collector's problem, we show it can be a logarithmic factor off of optimal. Other experiments show it can be even worse than this in realistic situations.
Our solution has two parts. We hash-and-partition the bad sources, adaptively finding a partition size that so that the bad sources in each partition fit into our small memory. That is, we hash each source, and put in a partition according to the last k bits. This breaks our N bad sources into groups of size (roughly) N/2^k, and if N/2^k is small enough so that the partition fits into memory, then (assuming we can avoid duplicates), we no longer have a memory problem. We just run through all the partitions, giving us our "Carousel".
To deal with duplicates, we do the obvious -- we use a Bloom filter (or equivalent structure) to avoid sending duplicates within each partition.
This hash-and-partition plus Bloom filter type framework seems like a potential generally useful trick; more details -- including both theoretical analysis and experiments -- are in the paper.
An interesting thing came up in the reviews. We pointed out that the "straw man" approach -- do nothing -- was quite bad. We mentioned that just using a Bloom filter -- without partitioning -- wouldn't really help, but then ignored that option. Apparently, this was a big concern for the reviewers; my understanding is that it almost "sunk" the paper. They wanted to see more details, including simulations, on this option. Luckily, it was still accepted, and in the final version, in response to the reviews, we've added some theory and simulations to prove our point. (The point is you'd need a really, really big Bloom filter -- too big for your memory -- to track all the sources, so eventually you have to clear the Bloom filter, which causes you to lose whatever it was going to gain you in the first place.) I'm not sure what the takeaway is there. The right outcome happened -- they should have accepted the paper, and we could add the appropriate stuff. But I'd have hated for what in my mind is a minor issue to have killed the paper. Perhaps we should have said more about it, although at some point space prevents you from providing details on every possible variation you could consider (even if you have actually considered them).
This seems like a good place to remind people about this book -- Algorithms for Next Generation Networks
Friday, March 05, 2010
SIGCOMM papers
I'm finally getting around to reading and reviewing. (First round reviews aren't due for at least a week!) And so far, by and large, the papers I'm getting are pretty terrible.
This generally seems to happen to me on the first round, but this year is extreme. My first several papers just don't belong at this conference. (Arguably, they don't belong at any conference...) There's some number of papers submitted at every conference that are just not serious submissions, and apparently I got more than my fair share on the first round.
This makes it harder to judge the other papers -- it's hard to calibrate when you start with a lot of junk. Many of my other papers are theoretically oriented, and I'm not too optimistic about them. There's room for theory papers at SIGCOMM, but I think the bar is, rightly, pretty high. When I read a theoretical paper for SIGCOMM, I look for one of two things. First, it could be the paper has a nice theoretical idea that's actually useful. The problem there is that it's incumbent on the paper to clearly demonstrate the utility, and most fall down in that regard. I quote the SIGCOMM call: "SIGCOMM is a highly selective conference where full papers typically report novel results firmly substantiated by experimentation, simulation, or analysis." A mathematical analysis alone generally does not count as a firm substantiation. [Such papers generally have a better chance at INFOCOM -- which I think is a good, and very important, thing! There needs to be an outlet for more theoretical networking work, and perhaps it's just better suited for a big conference. Many such papers will end up having minimal impact, but once in a while, a good idea gets built on and has a real impact.]
Second, it could be the paper really challenges our way of thinking, introducing a new framework that seems a clearly important guide for future work. Such papers are rare, but important. I seem to have a number of economics-networking papers that are very high-level, and I'm really trying to understand if any of them have that character. Again, because SIGCOMM is so selective, I think the bar is very high for such papers. I'm really looking for something that enhances our fundamental understanding of the network.
That's it for now. If you have a submission, don't let my comments make you antsy -- the meeting is still a long way away, and I'm quite sure I'm not reading your paper anyway.
Wednesday, March 03, 2010
Congratulations to David Johnson, Knuth Prize Winner
Tuesday, March 02, 2010
Teaching Bloom Filters
If you're teaching algorithms and data structures, do you students a favor, and sneak Bloom filters in one lecture.
Monday, March 01, 2010
Stuff in press
Bach, Chawla, and Umboh take our previous work on the hiring problem in new directions. Or, even, new dimensions: they also consider multidimensional variations of the problem. It's definitely a general problem with plenty of variations to consider; perhaps this paper will inspire further looks at the problem.
Friday, February 26, 2010
Conflicts of Interest, Yet Again
CONFLICT OF INTEREST GUIDELINES ============================================================================= A program committee member (including the chair of the committee) is considered to have a conflict of interest on a submission that has an author in any of the following categories: 1. the person themselves; 2. a past or current student or academic adviser; 3. a supervisor or employee in the same line of authority within the past five years; 4. a member of the same organization (e.g., company, university, government agency, etc.) within the past five years; 5. a co-author of a paper appearing in publication within the past five years; 6. someone with whom there has been a financial relationship (e.g., grants, contracts, consultancies, equity investments, stock options, etc.) within the past five years; 7. someone with whom acceptance or rejection would further the personal goals of the reviewer (e.g., a competitor); 8. a member of the same family or anyone considered a close personal friend; or 9. someone about whom, for whatever reason, their work cannot be evaluated objectively.
These guidelines are roughly the same (with minor variations) as what I've come to expect from other networking conferences. I feel I have to point out the remarkable difference between how conflicts are treated in the networking world and the theory world. In the theory community #1 is a standard conflict; #2 and #3 are also pretty standard although, in my experience, definitely not universally applied; and after that conflicts are generally, in my experience, up to the individual PC member to declare if they happen to feel like it.
There's been debate on this blog about the subject before, and I certainly don't mind there being more. I maintain that the theory community is far too lax in its handling of conflicts. We can certainly reasonably argue whether the true impact of conflicts in actual decisions in theory conferences is negligible or substantial -- a matter of appearance or a matter of substance. I can say that, in terms of appearance, people from the networking side (and other communities) are shocked by the lax approach adopted by the theory community.
Thanks to Muriel and Tim for allowing me to post from their document.
Thursday, February 25, 2010
STOC Budget Questions
Here's a bunch of questions that arise. I'm happy to hear input.
- Should all PC members' expenses for the PC meeting be paid for? That works out to, roughly, $80-100 on the registration per attendee. Hotels and airfare add up, and keep in mind the way the ACM forces us to do the budget you need to budget over 100% of the nominal cost to deal with contingencies.
In many other areas, it's assumed you'll pay your own way to the PC meeting. For the networking conferences I've PC'ed, they cover meals, and usually have a very nice dinner after the work is done. For the theory conferences I've helped manage, I've usually aimed to cover everyone's meals and hotel (though the dinner is less nice than for the networking conferences...), and to cover anyone who couldn't fund their own travel. That works out to more like $40 per attendee.
- Do we really need morning and afternoon coffee breaks? The afternoon coffee break every day adds something like $25 per.
- When you look at fixed cost, every student who attends is actually a loss, that has to be covered from elsewhere. Is this the right way to go? (I like to think that the corporate sponsorships, from Microsoft/Google/IBM/+others, should be first thought of as going to reduce the cost of student attendance, so I think this is still the way to go.)
- At what point do registration fees become a noticeable concern?
Tuesday, February 23, 2010
Guest Post from David Karger
------------
I wanted to post a comment on Mike's FOCS/STOC post, but fittingly for one of the dinosaurs he mentioned it was too big for the comment length limit. Mike's been kind enough to offer to post my comment as a guest post instead.
Mike's question is one I care about a lot. I still respect theory and do work in it, but as Michael says, much of my attention has been drawn into other areas. Many of them have absolutely nothing to do with theory (see last year's ethnographic study of people's use of pencil and paper for notetaking in TOIS 2009 or our AJAX-flavored interface for visualizing and navigating semistructured data in UIST 2009).
But always, some of my favorite projects are where theory provides the answers to problems that matter in other areas. We just published a paper in Nature Genetics that used some simple applications of max-flow to help biologists visualize the "important" influences in biological netwoks (we didn't need to find NP-hard integral solutions because the scientists wanted to see all the possibilities in the mix). Before that, we applied a beautiful JACM paper of Alon etc., on finding longish paths in a graph, to a problem in natural language processing---designing a procedure to figure out a best selection and ordering of words for a machine-generated summary of a machine-generated document---and published a paper at NAACL, a linguistics conference.. I still remember thinking, when I first read Alon et al., that it was one of the prettiest and cleverest ideas I'd seen in a while, but that it would never be useful for anything practical. Ironically, when my colleague Regina Barzilay outlined the language problem to me, it was exactly the Alon problem with no need for translation; my entire contribution was to know that a solution existed (and thus keep up theoreticians' reputation for smart enough to solve any problem instantly).
Other applications of theory to practice have required more work. Mike recently wrote about our paper showing how to design "network codes" for efficient multicast; the core insight of this work was to connect it to the beautiful results that we all study in randomized algorithms courses, on finding perfect matchings by placing random numbers in a graph's Tutte Matrix. Most substantially, our line of work that led to Danny Lewin and Tom Leighton's founding of Akamai begin with a study in STOC of some theoretical problems around handling flash crowds on the internet, and also generated a whole line of research on building robust and scalable peer-to-peer systems.
With the exception of the first paper on consistent hashing, none of this work has appeared in theory conferences. I think there are several reasons for this. Selfishly, for the author it is much more fun (and valuable in generating future research leads) to present the work at non-theory conferences. It holds the same attraction as tourism, going to strange new places and learning new things from the experience. There's also vanity---the allure of being an exotic theoretician among practitioners instead of one of a crowd of better theoreticians than you. Most important, if you want your work to have an impact on the applied areas, you have to publish in their conferences so they'll pay attention---they don't read STOC/FOCS.
But the second reason is more problematic. Many theoreticians would tell you (some have certainly told me) that the above papers are "not theory". That by dint of their having applications, they are no longer suitable for STOC/FOCS. That these papers have a different home, and FOCS/STOC should be reserved for "pure" (i.e. homeless) theory research. I've been on STOC/FOCS PCs that have rejected nice applications of theory on the grounds that the theory part was too elementary.
In part they are right. The biology and NLP papers I mentioned above did not prove any new theorems. But the omission of applications papers from STOC/FOCS means that theory community is failing to celebrate one of its greatest contributions! There's always been a divide in the theory community between those who are enamored of theory problems that help them understand the deep nature of the universe and computation (scientist/mathematicians) and those who see theory as a way of thinking about solving concrete computational problems that often emerge from other areas (engineers). I think that STOC/FOCS is making a mistake by focusing too much on the science to the exclusion of the engineering.
I would really like to see more "applied algorithms" papers appearing at STOC and FOCS. These are likely papers that have not made a major theoretical advance, but rather have synthesized our existing theory knowledge into a solution to someone's particular problem. These papers are just as important to see as the ones that advance theory; they represent one of the major justifications for doing theory in the first place.
I haven't mentioned SODA yet. That is a conference that was founded in part to attract these kinds of applications papers. But at the same time it was founded to attract more of the theoretical discrete math community, and the multitude of targets makes the outcome diffuse. Possibly as a result, SODA doesn't have the stature as STOC/FOCS; I'd like to see applied algorithms appearing at our flagship conferences.
Even if the STOC/FOCS community decides to do this, we still have to deal with what I said at the beginning, that the applied conferences are often more attractive for this kind of theory. You might say "fine, if that's what they want, there's no problem." But I think there is a problem: our community, and in particular our theory graduate students are not being exposed to this important branch of theoretical computer science.
The best solution I can think of is to allow repeat submission. That is, to let the paper appear first in the applied conference, then at STOC/FOCS. Almost by definition, these two venues will not have a lot of overlap, so I really don't see a downside to presenting such a paper at both of them. There are two ways to get this by the copyright police. The first is to accept the paper for presentation but publish only a reference to it in the proceedings. The second is to ask the author to write a new version of the paper aimed at a theory audience. Given the different audience, the paper is likely to be quite different.
Monday, February 15, 2010
FOCS/STOC : What's the Big Deal?
My point here is that many of the best theorists I know have, I would say, transcended FOCS/STOC. This does not mean there's not great stuff in FOCS/STOC; it just seems strange, given this, that these conferences are accorded such weight.
Perhaps, in fact, they're accorded less weight than I'm crediting them with these days. Certainly, the Innovations in Computer Science movement demonstrates some dissatisfaction with FOCS/STOC, and there are debates in subcommunities (SoCG, Crypto) about FOCS/STOC vs. the specialized conferences. It still seems to me, though, that FOCS/STOC is where most people would want their best theory results to appear, and it's still the lens through which fresh theory PhDs are viewed.
It seems to me that the theory community, as a whole, needs to think about FOCS/STOC/SODA and the other many conferences, and figure out what it wants them to me. FOCS and STOC haven't changed much over the years, and perhaps they've become just a bit too comfortable; the (theory) world around them has changed considerably, and it's not clear that they've adapted. Should they be the flagship conferences of theory, and if so, what does that mean, and how can they better fulfill that role? If they're not going to be the flagship conferences of theory -- which might be perfectly reasonable -- what is their role to be?
Saturday, February 13, 2010
News Roundup
One of those "denied-tenure-leads-to-shooting" incidents, this time in Alabama. (To be clear, the details about the reasons for the shooting are still, I think, unofficial.) I always feel a twinge when I hear a story like that... it makes me glad that Harvard's tenure process is extremely super-secret. A professor's job is, naturally, very safe, so stories of students or faculty losing it like this hit home. Oh, and I like how most news articles feel important to point out the shooter was "Harvard-trained".
Not-just-Harvard with budget woes. At least we seem to be reducing our red ink at a good pace.
Speaking of budgets, we have proposed increases in the NSF budget for the coming year. Although it seems to take advantage of it, you might want to start working in energy technologies. And there will be a new NSF director; I'm not sure how that affects us academics.
Any other news of note?
Thursday, February 11, 2010
FOCS 2010 Call for Papers is Up
Key points: The deadline is April 7. And, in a move that I approve of, there's no page limit on submissions -- instead, "Material other than the abstract, references and the first 10 pages may be considered as supplementary and will be read at the committee's discretion." Having just submitted papers to ICALP where we had to go through the "move things to an appendix" routine, I think this is the right way to go.
Tuesday, February 09, 2010
Recent Award for Network Coding
Sadly, it's also time to recall the passing of Ralf Kotter, who died of cancer just over a year ago. Communications Theory lost a brilliant mind and a great leader far, far too early.
Guest Post: Giorgos Zervas from WSDM, Part 3
Preparing to depart from New York, amidst rumors of a big snowstorm that never seemed to materialize, I was thinking of the other participants who didn't have the luxury of being just a four-hour drive away from home. Especially, those that had a long return flight ahead of them. After three days packed with talks, lunches, networking and even going out and enjoying what Brooklyn has to offer (a lot!), I am sure most people wouldn't have minded being teleported back. This makes me wonder: what are our main incentives for conference participation?
I am guessing one of them must be attending the actual talks. Which brings me to my next point... Sergej Sizov couldn't attend the conference and instead sent his presentation over: the usual slides accompanied by a video of him presenting the work. To be perfectly honest, I was rather negatively predisposed to the idea of a prerecorded presentation. And judging from the number of people in the auditorium I think more people may have thought the same. Yet, I was completely wrong. After 30 seconds or so, I was completely immersed and forgot that the presenter was a projection. The presentation itself was clear, finished on time and even got an applause at end - which was absolutely deserved (yet in the absence of the speaker reminded me of the awkward feeling I get when people clap in movie theaters.) I would say the only downside was that we had to skip the Q&A session. Technically, though I don't see why this couldn't have been arranged save for timezone considerations. So, if a taped delivery doesn't really compromise quality why don't we use this format more often and minimize travel? Could conference participation eventually evolve to a mixed model of in person attendance and participation over the web? I do realize the benefits of networking and meeting each other in person but do we really have to attend every single conference irrespective of cost and time issues?
Looking forward to WSDM'11 in Hong Kong!
Friday, February 05, 2010
Guest Post: Giorgos Zervas from WSDM, Part 2
Carlos Castillo did a fine job of presenting "An Optimization Framework for Query Recommendation" by Anagnostopoulos, Becchetti, Castillo and Gionis. My favorite part was using Cavafy and Machiavelli as presentation vehicles for two different utility functions they evaluated: the former aggregating utility along every step of a multi-step process, the latter ignoring the journey and solely caring about the value derived in the very last step. These utility functions were presented in the context of query reformulation, the query suggestions search engines provide users with to aid them in finding what they are looking for. I am not quite sure how they came up with this great metaphor but it may just be that the authors are Greek and Italian.
The second presentation I enjoyed was given by Alan Mislove. I think he nailed it by selecting just right level of abstraction for his talk. Not too much detail, but enough to maintain my interest and entice me to read their paper: "You Are Who You Know: Inferring User Profiles in Online Social Networks". The main idea here is that information that you may consider private and are unwilling to publish can potentially be inferred by information your friends reveal; not necessarily directly about you, but about themselves. Because of the homophily present in social networks what your friends say, can be telling about you. Hompophily was definitely word of the day today; it was featured in three different presentations. All in all I think this paper underlined some concerns anyone with an online presence should be having.
My only gripe so far has been the heat in the auditorium - I think, by the end of the day, it makes everyone feel more tired than they already are. But other than that WSDM has been very enjoyable so far.
PS: The Twitter feed disappeared during Thursday afternoon's sessions but was back this morning. I guess people must be enjoying it!
Guest Post: Giorgos Zervas from WSDM
-----------------------
Greetings from Brooklyn.
I am here attending WSDM 2010 where on Saturday I will be presenting the work we did with John & Michael on Adaptive Weighing Designs for Keyword Value Computation. Michael asked for my grad-student perspective on the conference and I happily obliged as he offered me a decent revenue-share deal on any book sales resulting from this post.
On a more serious note, while being here, my primary concern is presenting our work in the best possible light and hopefully getting some people to read the actual paper. A secondary, but equally important concern, is that of networking. My impression is that for most conference participants time is a very scarce resource. Of course this is not limited to WSDM. The sight of a speaker surrounded after his or her talk by a bunch of people - just like myself - intimidates me and makes me feel that by adding myself to the pool, I am becoming an additional burden to someone who might have better things to do than listen to my 30 second blurb (even though personally I'd be flattered, so please surround me after my talk!). Ideally, I'd prefer interaction to be more organic and there are certainly some good opportunities for that. My question to you is: do you prefer some ways of being approached over others? How do you respond to cold introductions? Any advice on how grad-students should network at conferences? And if you are student: what do you find works for you in terms of introducing yourself to others?
On a different note, an interesting feature of WSDM has been the live Twitter feed projected behind the speaker, next to the actual presentation slides. Even though some of the tweets are insightful and they make for great conversation starters over lunch, I find the projection rather distracting. I reckon that, for 20 minutes, focus should be on the speaker and the tweets are almost impossible to ignore. Some of them are also quite repetitive (videolectures.net anyone?) Those wishing to follow the Twitter stream could always do so from their laptops and phones. What are you thoughts? Do you think this backchannel adds to the discourse? (I should point out that the WSDM community definitely seems mature enough to avoid mishaps like this.)
Finally, on the research front, and even though we are still on the first day of the conference, I've had the opportunity to attend some great talks. In particular Soumen Chakrabarti gave this morning's keynote and I found his vision of extracting structure from the unstructured web fascinating. A few papers that have grabbed my attention (in no particular order) are "Automatic Generation of Bid Phrases for Online Advertising" (Ravi et al.), "SBotMiner: Large Scale Search Bot Detection" (Yu et al.) and "Evolution of Two-Sided Markets" (Kumar et al.); I am looking forward to these and the rest of the presentations. What have been your personal favorites?
Thursday, February 04, 2010
Admissions Handling
For graduate admissions, we've moved to an all-electronic system; the applications are all online. The system is actually a complete pain to use. Does that surprise anyone? (Some of us get our admins to download everything for us, so we don't have to use log in and navigate the system when we need to look at an application, or spend an hour or more ourselves downloading files in a system that wasn't set up to download selected files in a straightforward way.) While I'm absolutely, positively, completely sure that no candidate's confidential information has ever, ever been compromised, or ever will (I believe I've now covered myself and Harvard legally), it seems like a privacy-risk nightmare with all the applications secured by a password that has to get distributed to all the faculty. Still, with all that, it seems slightly better than the paper folder system we had before, where it seemed impossible to track which professor had which folder, never mind actually arranging for folders to be transferred among multiple faculty in a timely manner.
For undergraduate admissions, I'm asked to look at folders -- usually, I'm being used to check that Johnny or Jane's science fair project actually has some interesting science in it, or similarly vouch for math/science talent, but they seem to appreciate if I make other comments as well. It's all paper. An actual person drops folders (a few a week) off to my admin; I type up comments and my admin prints them out, puts them into the folder, and calls to get the folders picked up. Apparently, it's unusual that I type my comments; the admissions officers write their comments by hand. (Often, I can't read them, and my handwriting is worse than theirs.) I've never lost a folder, but I do hope they have back-ups in the home office just in case. (It looks to me like I'm getting the originals; I've never asked. I just assume they can't give the folder out to faculty without keeping a copy of everything.)
Both systems seem flawed, but both also seem designed to fit the way the decision-process is made. I actually like the paper system, even though it clearly requires a lot of people-hours doing background tasks like getting folders from here to there. It definitely reduces the time and effort I have to put in to review the applications -- which ostensibly should be the goal of the system, since faculty time is (ostensibly) valuable.
Tuesday, February 02, 2010
Does Class Size Matter?
Why should I, or any faculty member for that matter, care about class size? For junior faculty, at least, there's a clear answer. A tenure case without some teaching of significantly sized classes is one with a weakness, opening the way to the arguments that the faculty member in question is working in an area of little interest (since nobody wants to take their classes in their area) or is providing insufficient service to the department (since they haven't taken on a large core course).
For senior faculty, I'm not so clear. While class sizes are listed in our annual review, I've heard no mention that they're considered of any particular importance -- or even of non-zero importance -- in determining annual raises. Indeed, I can't think of any direct benefit to me personally for teaching a large undergraduate class as opposed to a small one.* One might want to take on a big class to support one's department, as ostensibly money (and positions) should, in some way, follow students at the departmental level. I'd like to think that's how it works at many places; however, recent conversations with some higher-ups suggest that that connection is fairly tenuous for SEAS. Harvard's system in that respect seems to be broken. (If it wasn't, I think we'd be further ahead in our hiring in CS.)
So why should I care about my class size? Primarily, I suppose, personal pride. I take satisfaction in teaching students; the more qualified students, the better. (Not the more students the better, though; the more qualified students, the better...) Indeed, I've done the math, and while I'm quite sure Harvard does not calculate things this way, in my mathematical model I'm earning what Harvard's paying based solely on the number of students I teach. That helps me sleep at night.
Overall, however, this seems like an area where the incentive structure doesn't seem set up right. I can understand that class size isn't an end in itself; indeed, I can understand that part of the mission of the University is to preserve knowledge in areas that might be of narrow interest. (The Sanskrit, Slavic, Turkish, and Yiddish courses, for instance, have remarkably low numbers.) But it seems naive to think that size doesn't matter **, so it's slightly disturbing that when I think in terms of incentives, I'm ending up wondering why I should care about my class size at all.
* I do see a potential direct benefit for having my large graduate project class; some student projects can get turned into papers, and often students have me take part in turning their project into a paper, so I may get some research benefit from having a large graduate class. It's not clear that's a big benefit, but at least it's demonstrable.
** Yes, we all knew that was coming before the end of the post....
Sunday, January 31, 2010
Justifying Growth : We Need Better PR...
I wandered into lunch the other day and saw some other faculty from SEAS -- but from well outside computer science -- already eating, so I joined them. In our friendly discussions, they asked about our growth plan, and asked me to justify it further to them. One point I brought up was that we did a lot of "service" to the rest of the university. Sure, they said, they knew about the very large intro programming class, but that was just one class. What else?
I listed off several of our other classes that they didn't seem really aware of -- our course for non-majors CS 1 (Great Ideas in Computer Science) and our Gen Ed course Bits, our more advanced programming classes CS 51 and CS 61, our new interdisciplinary visualization course CS 171 and our new course on design of usable interactive systems CS179, and probably a few others. I then mentioned that even our theory courses (introduction to complexity, and introduction to algorithms and data structures) were attracting a lot of non-majors, and pointed out that they each had 80 students last year. (We have about 30-40 or so CS majors a year right now.)
Their jaws literally dropped. One of them asked me, multiple times, how there could be 80 people at Harvard who were interested in taking Algorithms. (While I, of course, always wonder why it's so few.) I still have doubts that they believed me; I think I ought to get something official-looking from the registrar and send it to them.
Now, admittedly, last year's class was big -- my class this year looks to be a more normal about 50 or so. But I was still surprised by their surprise. I check out the class sizes around SEAS to get a feel for what's going on every year (we usually get an e-mail with course counts). Also, I'm pretty sure I mention my class size fairly often to other SEAS faculty when the opportunity arises, but apparently less often than I think. I was left with the feeling that I, and maybe the rest of the CS faculty, needed to engage the other SEAS faculty a bit more and let them know more about what we're doing. In particular, as another faculty member said to me, "When they talk about a really big class, they mean 40 students. When we talk about a really big class, we mean over 100."
I'm not exactly a shy, retiring type. (I mean, c'mon, I blog.) But I'll be upping my efforts to make sure others at Harvard have a better idea of what we're doing, especially in terms of teaching our undergraduates.
Friday, January 29, 2010
Paper updates
Also now up on the arxiv is a preprint (submitted) by myself and Raissa D'Souza on Local cluster aggregation models of explosive percolation. We were looking at variations of the Achlioptas process. (I've promised myself I'll stop making fun of Dimitris for his keen ability to get a process named after himself -- sometime around 2020.) In the basic Achlioptas process you start with an empty graph of n nodes, and at each step you choose 2 edges at random, and add one to the graph according to some deterministic rule, such as add the edge that minimizes the size of the resulting merged component. Usually there is an associated goal, such as to delay (or speed up) the emergence of a giant component in the graph. As you might expect, the power of choosing an edge allows one to potentially delay or speed up the emergence of a giant component substantially. Our paper looks at what seem to be a simpler class of similar processes, the most basic of which is that at each step you choose 2 edges at random that are adjacent to some randomly chosen vertex v. That is, we look at local variations, where the choice can be thought of as the vertex v choosing which of 2 random edges to add to the graph. It doesn't seem these local variations have been analyzed previously. We look at differential equations that describe the underlying process and empirically examine whether such processes have discontinuous phase transitions, like the original Achlioptas process appears to have. We're promoting the idea that these "local variations" may prove simpler to analyze fully and rigorously than the original variations, for which there still remain many open questions.
Addendum: Just in case there's any confusion, I should add that Dimitris in no way named the Achlioptas process after himself. (I was just teasing him.) He was the first to suggest and promote the study of the process and when others wrote the first papers about it they nicely credited him by calling it an Achlioptas process. The name has rightfully stuck.
Tuesday, January 26, 2010
Teaching, Day One
Last year there were about 80 students in the class, a remarkable reversal after several years of decline -- the year before that there were fewer than 40. Of course, this was not successful for everyone, so even though our intro programming class has continued to grow, I'm expecting fewer students this year. The fall theory class (our intro complexity class) had about 60 students, and usually my numbers are a few less than that. They had about 80 students last year too. So my expectation for the final course size is in the (mid) 50s.
When I got to class I wondered if the room assignment hadn't been posted -- only about 20 people showed up. But lots of people wandered in a few minutes late, and then more and more as the class went on. The joys of Harvard's shopping period (though I can't recall ever quite such a number of late arrivals). I'm pretty sure at least 60 showed up at some point during the class; my initial estimate may be pretty close.
What's very odd this year is that on day one there's 48 people signed up for the Distance Education version of the class through the Extension School. (Still time to sign up!) That's definitely more than normal. A big fraction of those students are likely to drop out -- many quickly figure out the class is more than they can handle -- but that's still a much larger starting number than usual.
Hopefully, by next week, I'll have more accurate numbers for both versions of the class.
Lecture went fine. Very well, in fact, in that a number of students raised their hands in response to my questions, and they had, as a whole, very good answers and insights. Optimistically, I'm looking forward to an above-average class this year.
Sunday, January 24, 2010
On Formatting
But isn't it time that we define a single standard for conference paper formatting that everyone uses?
(With such a system, conferences could still vary the length of the paper they accepted -- but the style file would be the same.)
I think that would be a great step. In the same spirit, I'd also like to see less focus on page limits -- both for submissions and for the final version. (Ostensibly, the final version should be closely related to the actual submission -- so if you're going to relax page limits for the final version, it seems best to do so for the submission version as well.) I think the recent change for SODA -- allowing up to 20 page conference papers -- is a great idea.
I understand that, in some cases (where printing is involved), some sort of page limit may be needed -- although with less and less printing of full proceedings, it's less clear that this is an important consideration. I also believe that page limits can force people to improve their writing -- bad writing and excessively wordy writing generally go together. (If we remove page limits, we'll have to tell people in reviews more frequently when they need to cut things down or even out in their papers.)
But the effort spent on meeting arbitrary page limits -- just the time spent on formatting -- seems silly at this point. And often it forces you to cut content. I can't count the number of times I've had a review say, "You should have included this...", where my response would be, "We did include this, but had to cut it to make it fit into the page limit..." (Yes, you can create multiple versions, and post the full versions online -- except for double-blind conferences, like SIGCOMM, of course -- but that involves yet further overhead, when you have to update to deal with reviewer comments, and...) And if your paper is rejected from one conference, to submit to another, you have to re-format and create another version to meet their arbitrary format and page limit standards.
In some ways, I'm glad that Matt is out there, talking about the papers being rejected for formatting. And I hope it starts happening to more people, more often. Because I worry that the only way for change to happen is for bad things to start happening to good papers so that people will start to realize that the current system is just bizarre and broken, and a revolution needs to occur. Perhaps I'm wrong, and slow evolutionary change -- with 20 page papers a SODA being an example -- will lead us the right way. If so, I encourage PC chairs to experiment with flexibility in their formatting requirements.
Tuesday, January 19, 2010
Housework
Monday, January 18, 2010
New TCS postdocs/jobs website at CCI
Thursday, January 14, 2010
Letters and Rights
I simply prefer my confidential letters to be confidential. I understand that waiving their rights is not any sort of guarantee that they won't see the letter, but I think the understanding is important: I'm writing an evaluation of them to my colleagues, and that information is, as a matter of default policy, not meant for their eyes, so I can be forthright in expressing my opinions.
Wednesday, January 13, 2010
Algorithms and Data Structures : Course Goals
I have a somewhat different take on the problem than Richard, though. He ends up creating a list of things (definitions and theorems) he thinks student should know after his class. I admit I think a bit more in terms of high-level concepts and skills I want students to get out of the course. (I think both approaches can work.) My list includes:
1. Students should understand the basics of algorithm/data structure language and notation; in particular: order notation, what it means (and doesn't mean!), and how to calculate running times of algorithms.
2. Students should learn how to model a problem through proper abstraction. In particular, they should learn how to set problems into the language of graphs, linear programs, etc.
3. Students should learn basic algorithmic techniques: greedy, divide and conquer, dynamic programming, linear programming.
4. Students should learn what is required for a formal and rigorous argument (even if they are not so good at writing one themselves). They should learn how to argue/prove correctness of algorithms and data structures in the contexts above.
5. Students should clearly understand the power of reductions -- not just to show problems are hard, but to show problems are easy, by using the solution of one problem to solve another.
6. Students should obtain some insight into the power of randomness in algorithmic thinking.
7. Students should practice translating algorithmic ideas into working code. It is hoped they will learn the importance of thinking before coding (choosing the right approach), and that issues such as memory requirements and running time can and should be reasoned about before implementation.
8. Students should learn that correctness, like computation time and memory, is just a "resource" that can be traded off with the others, by learning the notion of an approximation algorithm.
9. Students should learn the basic of heuristics and heuristic techniques, so that if asked in practice to solve an NP-hard problem (in the sense of coming up with a very good but not optimal solution), they have some idea of what approaches might give them reasonable answers.
10. Students should see some really cool, unusual, interesting algorithms and data structures, and be forced to work on some really hard, challenging problems, to challenge their thinking and demonstrate the power of theory.
There's probably more I could add, but those seem like ten clear objectives I have for the students who take my class. Perhaps in another post I can elaborate on how I try to see those objective are met.
Like Richard, I should end the post by asking the important open questions:
What have I left out? What should I leave out? What do you think?
Tuesday, January 12, 2010
New Harry Lewis Opinion at HuffPo
Larry Summers, Robert Rubin : Will The Harvard Shadow Elite Bankrupt The University And The Country?
gives some idea of where he's coming from...
Saturday, January 09, 2010
Visiting Dartmouth for a Research Symposium
They've done a very nice job in organizing it - they seem to have corralled most students and faculty into coming, and they have brought in huge quantities of food (a known motivator for student attendance). [Note to self: Dartmouth CS apparently gets its breakfast muffins from "Lou's". Those things are amazingly good.] The presentations are going well and I'm sure it's a useful experience for students. So I'm trying to learn from this how we might do something similar at Harvard.
We used to have something like this at Harvard -- we had an "Industrial Partners" day where we'd try to bring people from labs/companies in and do posters and presentations (including some by graduate students) in a similar fashion. At some point, it fell below critical mass, but we're perennially thinking of how to bring something like it back. You do have to get people to commit to it, and to find time for it -- notice Dartmouth has chosen to hold it on a Saturday, and at a time where the campus is otherwise pretty dead. Maybe that's the approach that would work for us. I've found it's hard to schedule anything on a regular M-F time slot with a group of faculty, even for a hour, never mind a whole day.
Another issue with scheduling a research symposium is that we, already, suffer from talk overload. We have our own colloquium; the various groups (theory, AI, systems) run their own seminars or lunch-talks; there's plenty of other talks around Harvard run by various departments or organizations; and there's a seemingly endless number of talks nearby (at MIT, Microsoft NE, and other places). The problem about having so many talks is you start to lose interest in organizing more talks, and arguably the graduate students benefit at least as much by presenting in their own group seminars. But visiting Dartmouth today has given me some further evidence that the good that can come out of a symposium day like this might make it worthwhile. Perhaps the key is to let the graduate students plan and run the day, so that they own it, and feel responsible for making it a success.
How many of you have a similar research symposium day, how does it work, and what do you think of it?