Saturday, 29 March 2008

Wikipedia's downstream traffic

We've been hearing for a while about where Wikipedia's traffic comes from, but here are some new stats from Heather Hopkins at Hitwise on where traffic goes to after visiting Wikipedia. Hopkins had produced some similar stats back in October 2006, and it's interesting to compare the results.

Wikipedia gets plenty of traffic from Google (consistently around half) and indeed other search engines, but what's interesting is that nearly one in ten users go back to Google after visiting Wikipedia, making it the number one downstream destination. Yahoo! is also a popular post-Wikipedia destination.

It was nice to see that Wiktionary and the Wikimedia Commons both make it into the top twenty sites visited by users leaving Wikipedia.

Hopkins also presents a graph illustrating destinations broken down by Hitwise's categories. More than a third of outbound traffic is to sites in the "computers and internet" category, and around a fifth to sites in the "entertainment" category, which probably ties in with the demographics of Wikipedia readers, and the general popularity of pop culture, internet and computing articles on Wikipedia.

Hopkins makes another interesting point on the categories, that large portions of the traffic in each category are to "authority" sites:

"Among Entertainment websites, IMDB and YouTube are authorities. Among Shopping and Classifieds it's Amazon and eBay. Among Music websites it's All Music Guide For Sports it's ESPN. For Finance it's Yahoo! Finance. For Health & Medical it's WebMD and United States National Library of Medicine."

Similarly, Doug Caverly at WebProNews states that the substantial proportion of traffic returning to search engines after visiting Wikipedia "probably indicates that folks are continuing their research elsewhere", and this ties in well with Hopkins' observation about the strong representation of reference sites.

All of this suggests that Wikipedia is being used the way that it is really meant to be used: as a first reference, as a starting point for further research.

Monday, 17 March 2008

Protection and pageviews

Henrik's traffic statistics viewer, a visual interface to the raw data gathered by Domas Mituzas' wikistats page view counter, has generated plenty of interest among the Wikimedia community recently. Last week Kelly Martin, discussing the list of most viewed pages, wondered how many page views are of protected content; that thought piqued my interest, so I decided to dust off the old database and calculator and try to put a number to that question.

The data comes from the most viewed articles list covering the period from 1 February 2008 to 23 February 2008. I've used that data, and data on protection histories from the English Wikipedia site, to come up with some stats on page protection and page views. There are some limitations: I don't have gigabytes of bandwidth available, so some of the stats (on page views in particular) are estimates, and protection logs turn out to be pretty difficult to parse, so I've focused on collecting duration information rather than information on the type of protection (full protection, semi-protection etc). Maybe that could be the focus of a future study.

There were 9956 pages in the most viewed list for February 1 to February 23 2008. Excluding special pages, images and non-content pages, there were 9674 content pages (articles and portals) in the list. Interestingly, only 3617 of these pages have ever been protected, although each page that has been protected at least once has, on average, been under protection nearly three times.

Protection statistics

Only 1223 (12.6%, about an eighth) of the pages were edit protected at some point during the sample period, 902 of those for the entire period (a further 92 were move protected only at some point, 69 of those for the entire period). Each page that had some period of protection was protected for, on average, 82.9% of the time (just under 20 days), though if the pages protected for the whole period are excluded, the average period spent protected was only 34.8% of the time (just over eight days).

The following graph shows the distribution of the portion of the sample period that pages spent protected, rounded down to the nearest ten percent:


The shortest period of protection during the period was for Vicki Iseman, protected on 21 February by Stifle, who thought better of it and unprotected just 38 seconds later.

Among the most viewed list for February, the page that has been protected the longest is Swastika, which has been move protected continuously since 1 May 2005 (more than 1050 days). The page that has been edit protected the longest is Marilyn Manson (band), which has been semi-protected since 5 January 2006 (more than 800 days).

Interestingly, the average length of a period of edit protection across these articles (through their entire history) is around 46 days and 16 hours, whereas the average length of a period of move protection is lower, at 41 days 14 hours. I had expected the average bout of move protection to last longer, although almost all edit protections do include move protections.

The next graph shows the distribution of protection lengths across the history of these pages, for periods of protection up to 100 days in length (the full graph goes up to just over 800 days):

Note the large spikes in the distribution at seven and fourteen days, the smaller spike at twenty-one days and the bump from twenty-eight to thirty-one days, corresponding to protections of four weeks or one month duration (MediaWiki uses calendar months, so one month's protection starting January will be 31 days long, whereas one month's protection starting September will be 30 days long).

The final graph shows the average length of protection periods (orange) and the number of protection periods applied (green) in each month, over the last four years:

At least on these generally popular articles, protection got really popular towards the end of 2006 into the beginning of 2007, and again a year later. However, it seems that protection lengths peaked around the middle of 2007 and have been in decline since then.

Protection and pageviews

What really matters here though is the pageviews. The 9674 content pages in the most viewed list were viewed a total of 805,569,269 times over the relevant period. The 1223 pages that were edit protected for at least part of the period were viewed a total of 270,057,550 times (33.5%), with approximately 247 million of these pageviews coming while the pages were protected.

This is a really substantial number of pageviews, however, this number includes the Main Page, which alone accounts for more than 114 million of those pageviews. Leaving the Main Page out of the equation gives a healthier figure of around 133 million views to protected pages during the relevant period (and remember, this is only counting pages on the most viewed list).

Conclusions

Although only one in eight of the pages in the most viewed list were protected at some point during the relevant period, they tended to be higher-profile ones, accounting for one third of the page views. The pages that were protected at some point tended to be protected alot of the time, three-quarters of them for the entire sample period. This certainly fits with what many people have already suspected, that a small pool of high-profile articles attract plenty of attention in the form of both page protection and page views.

It will be interesting to do some more analysis on the history of page protection. Based on just this small sample, it seems that average protection lengths are trending downwards, which could well be something to do with the advent of timed protection. Hopefully I'll have some more insights to come.

Sunday, 16 March 2008

Today's lesson from social media

So much for claims that Jimmy Wales uses his influence to alter content for his friends: according to Facebook's Compare People application, Jimmy may be the best listener, the best scientist and the most fun to hang out with for a day, but he's nowhere to be seen in the list of people most likely to do a favour for me.

Wednesday, 20 February 2008

Beware corners

I'm sure that everyone who follows the news around Wikipedia will be aware of the latest controversy to gain attention in the media, namely the dispute about the inclusion of certain images in Wikipedia's article on Muhammad. Much of the external attention has focused on an online petition that calls for the removal of the images which, at the time of writing, has more than 200,000 signatures.

The debate so far has been understandably robust. Unfortunately, issues like these tend to harden positions, and push people towards the extremes. Consider a recent example: the seventeen Danish newspapers who, in the wake of the arrest of three men suspected of planning to assassinate Kurt Westergaard, author of one of the cartoons at the heart of the Jyllands-Posten Muhammad cartoons controversy, republished the cartoons in retaliation.

Likewise, positions are being hardened in this debate among both supporters and opponents of the images. The relevant talk pages are remarkably free of comments (from either side) even contemplating compromise. The Foundation is receiving emails on the one hand giving ultimatums that the images be removed, and on the other exhorting the Foundation not to "give in" to "these muslims [sic]".

Retreating into corners like this is contrary to the ethos of Wikipedia, which operates on open discussion in pursuit of neutrality. So just as the supporters of the images are asking opponents to challenge their assumptions, so too should the supporters be prepared to challenge their own.

The first assumption that should be questioned is that the images are automatically of encyclopaedic value. Images have little value in an encyclopaedia unless used in a relevant context and given sufficient explanation. Take this image, for example. An interesting image, but unless it is explained that it appears in Rashid al-Din's 14th century history Jami al-Tawarikh, and the observation made that it is thought to be the earliest surviving depiction of Muhammad, it lacks its true significance. We have a whole article on depictions of Muhammad. While some of the images in it are discussed directly, many are merely presented in a gallery, without much text to indicate their importance or relevance.

The second assumption worth revisiting is the assumption that, since the images were created by Muslim artists, then there are no neutral point of view problems. This view overlooks the fact that there are many different traditions within Islam, not only religious ones but artistic ones also. The Almohads, for example, with their Berber and eventual Spanish influences, had vastly different cultural and artistic influences than the Mongol, Turkic and Persian influenced Timurids. The Fatimids of Mediterranean Africa had different influences again from the Kurdish Ayyubids.

The Commons gallery for Muhammad contains an abundance of medieval Persian and Ottoman depictions, a small handful of Western depictions, but only one calligraphic depiction, and no architectural ones. Calligraphy is extremely significant in Islamic art, given the primacy of classical Arabic as a liturgical language in all Islamic traditions. It's worth considering why there is such an over-representation of Persian and Ottoman works, and such a dearth of works from other Islamic traditions. It's worth considering for a moment whether the Western preference for natural representations, as opposed to the abstract representations preferred in most Islamic traditions, has informed the predominance of physical depictions of Muhammad in the English Wikipedia and on Commons.

These images should not be removed altogether; many come from historically significant works, and represent a significant artistic tradition. But the images - as with any other content on Wikipedia - ought to be used in appropriate and expected contexts, and ought not be used exclusively or primarily to illustrate these articles, but should be accompanied by images representative of other traditions.

Most of all, discussions on these questions should proceed openly and freely, and all participants should make an effort to question their assumptions, and move away from their corners.

Wednesday, 19 December 2007

Knol worries

Last week Google announced an invite-only trial of a new tool called Knol (their name for a "unit of knowledge") to allow people to write an information page on a subject which can then be rated, reviewed or commented on by others. The central idea, as Google's VP of Engineering Udi Manber put it, is authorship:

"Books have authors' names right on the cover, news articles have bylines, scientific articles always have authors -- but somehow the web evolved without a strong standard to keep authors names highlighted. We believe that knowing who wrote what will significantly help users make better use of web content."

You can see a sample knol here (isn't everyone just itching to edit out that spelling mistake in the first sentence?).

Most of the press coverage of Knol is positing it as a competitor to Wikipedia, but is it really? We won't know what Knol will really be like until it is open to the public (presuming it makes it out of private beta), but from the looks of things it differs markedly from Wikipedia in all of the most important ways that Wikipedia is unique.

Firstly, knols won't be collaboratively written: Google says that the Knol platform will include "strong community tools", enabling the general unwashed to submit changes to knols (they use the name for individual articles too) as well as review, rate and comment on them, but ultimately the content of knols will be controlled by their original authors. Obviously, this is different from Wikipedia's collaborative wiki editing model under which no-one owns articles.

Secondly, there will likely be multiple knols on any given subject: as they say in the Knol announcement, it will be Google's job to appropriately rank knols in search results. Presumably they'll make use of the rating and reviewing tools in the platform as well as standard metrics like PageRank to try to work out which knol really is the most authoritative on a subject. Again, this is clearly different from Wikipedia, with its single-voice, neutral point of view system.

Thirdly, knols will not necessarily be free-content: while the sample knol mentioned above has a CC-BY 3.0 licence displayed on it, there are no indications that such licencing will be required, and Knol's "author control" vibe probably indicates that each author will get to choose the licence for their knols.

Kevin Newcomb at Search Engine Watch thinks a better comparison for Knol is Squidoo, and to an extent Mahalo, "since it allows users to build authority and sign their work [and aims] to build content pages that rank highly in search engines." Danny Sullivan at Search Engine Land also draws the comparison between Knol and Squidoo, and suggests that Knol is more likely an attempt by Google to carve out a niche of its own in the 'knowledge aggregation' industry rather than an effort to compete directly with any of the projects in the field.

However, it's perhaps best to think of the Knol proposal less as a project and more as a platform: Rafe Needleman at Webware compares Knol with Google's existing text publishing platform, Blogger, though with "Digg-like elements". Knol authors will build reputation, like blog authors (though Knol will be more about discrete articles rather than a stream of them), and users will rate and review competing knols in much the way that Digg and similar link-sharing sites operate. I think this is the best comparison, and fits well with the strong focus on individuality and authorship, and on Google's planned hands-off approach, in the Knol announcement. There's certainly a niche available for this kind of publishing.

So, presuming Knol goes public one day, it may well garner a significant slice of search results and a place in the knowledge business, but with its author-driven multiple-voice model and basis as essentially a publishing platform, it is more likely to be a complement to Wikipedia than a competitor.

Wednesday, 12 December 2007

Mr Wales Goes to Washington

Jimmy Wales testified before the United States Senate Committee on Homeland Security and Governmental Affairs on Tuesday (Washington time) on collaborative technologies generally and the Wikimedia projects specifically, and how that relates to e-Government initiatives in the United States.

Among other things, Jimmy discussed the use of both internal and public-facing wikis for governmental communication, explained semi-protection to Joe Lieberman, and outlined the benefits of systems that are designed to be open rather than closed. Also testifying were Karen Evans from the US government's Office of Management and Budget, John Needham from Google, and Ari Schwartz from the Center for Democracy and Technology.

You can view video of the hearing on Youtube or download it here in Real format (be aware that it's two hours long), see Jimmy's prepared testimony here (PDF) or see other statements here.

Saturday, 3 November 2007

One year of Citizendium

Citizendium has turned one year old (at least, it's been one year since its initial pilot release) and the Citizendium community is looking back on what they have done so far, and looking forward to what they aim to achieve in the next year and beyond.

Citizendium founder Larry Sanger has posted a "one year on" status update on the project. In it he addresses some "myths" about the project, many of which relate to its expert-led content generation model, and many of which focus on the number of articles that the project has produced so far. Sanger is particularly strident about rejecting the suggestion that Citizendium is just another Nupedia, and that Citizendium is a poor competitor with Wikipedia.

While it's true that Citizendium's model does differ from that of Nupedia, one key similarity is that is the role of experts to have the final say in approving articles. Sanger is keen to point out that Citizendium has more than 3,200 "live" articles, but a "live" article can be started by anyone, and includes articles imported from Wikipedia to which at least three "significant" changes have been made.

Given the expert-led content generation process by which it seeks to differentiate itself, the real measure of Citizendium has to be its output of "approved" articles, the ones which have been approved and locked off by the experts. After a year Citizendium has only 39 approved articles, a little more than the 24 approved articles which Nupedia generated through its existence.

One of the key issues facing Citizendium has been its rate of contributor growth. Sanger predicts that eventually the project will reach that critical mass whereupon it begins to experience rapidly accelerating, even exponential growth in contributors as awareness of Citizendium spreads. But that's not looking likely at the moment: while user activity is up, overall user numbers have been fairly constant for some time now, suggesting that while the users who are there are getting more and more involved, the project is not attracting many new editors.

I think that if Citizendium is to succeed it needs to communicate a clearer identity to the public at large. I suspect that the common perception of Citizendium is that it is merely an expert-run competitor to Wikipedia, and the truth of this aside, the comparison between the two projects seems to be something of a chip on the shoulder of at least part of the community at Citizendium. Consider the recent call for essays on which licence Citizendium should be using for its content; this essay opposes using the GFDL "both for its own sake and because Wikipedia uses it", and Mike Johnson references the debate within the community as to whether or not Citizendium's content should be compatible with Wikipedia. Go see the "Citizendium and Wikipedia" board on their forums for an idea of what I'm talking about.

Maybe Citizendium should look to differentiate itself in more ways than simply having experts at the top of the editorial food chain. Topic selection and content style are obvious areas (see also this discussion), but there are many other aspects of content policy which on Wikipedia have been influenced strongly by the fact that the content generation model is open and not formally reviewed; analysis and synthesis, for example, which are not permitted on Wikipedia, are ideally suited to a collaborative model which also includes expert oversight.

There is promise in the idea of Citizendium, but after seeing the results of one year of existence, I am not convinced that the predicted "coming explosion of growth" is inevitable.