Waxy.org
 
Waxy.org is the sandbox of Andy Baio, an independent journalist and programmer living in Portland, Oregon. I created Upcoming.org and some other stuff too.

Contact Me: log@waxy.org or waxpancake on AIM

Tracking Twitter's Message Growth

Posted Mar 15, 2007

By now, everyone knows that Twitter exploded at SXSW and everybody's seen the Alexa charts. But this is mostly a mobile app, so pageview traffic is only part of the story. How fast is Twitter really growing?

I decided to find out by using Twitter's founder Evan Williams himself, albeit indirectly. Since Ev's Twitter history goes from message #28 in March 2006 to #8,281,991 about three hours ago, it's a convenient snapshot of Twitter's growth since it began. Update: The data from November 2006 to present is faulty, since they apparently switched to non-sequential IDs. More information below.

I threw it all in Excel and charted the sequential IDs and dates for each of Ev's 1,226 messages. The moment SXSW started, Twitter's growth curve changed radically and hasn't slowed down. (The huge orange bar for March is only half the month.) But more interesting to me are two other dates: November 23, when Twitter's growth rate sped up drastically and then on February 5, when the rate seriously slowed. Are there problems with my data? If so, I can't find it. If you have any sense of what triggered those changes, please comment and let me know.

To help with your search, and any other visualization, I've posted the full Excel spreadsheet with inline charts and tables. Enjoy!

Note! The last three are cumulative charts, not month-to-month growth numbers. That means there were not 8 million messages sent this month, but 8 million total since Twitter started.

Twitter Growth

Update: Jason and I just discovered that the IDs since November 2006 have not been sequential, rendering these charts useless. The jumps in activity were largely artificial. Jason has more information.

27 Comments (Add Yours)

Mar 15, 2007
4:49 PM  
/pd wrote:

nice.. the long tail in motion :)-

thanks for sharing !!


Mar 15, 2007
5:14 PM  
Myles wrote:

The slow down is probably due to Twitter hitting the performance wall. I know I had a lot of trouble posting to and receiving messages from Twitter during that period -- so I lowered my usage of it until performance improved.


Mar 15, 2007
5:17 PM  
Andy Baio wrote:

Hmm, it looks like November 23 was Thanksgiving. (Though I don't understand why that would lead to a change in the growth rate.)


Mar 15, 2007
5:54 PM  
Michael Specht wrote:

Maybe late November was when Twitter first starter to appear outside of the really early adopters, for example Scoble signed up 20 Nov?


Mar 15, 2007
6:15 PM  
Jeb wrote:

I second the performance wall hypothesis. That's when I started implementing my "connect twitter to arbitrary desktop apps" project (mostly just to practice Proto/VBA stuff, I don't know why anyone would want every iTunes track to be a twitter message) Twitter was reaaaaaallly slow. I'd played with it a long time before then and it was snappy. I know in that couple weeks when I was playing with it, I personally would have sent several times more messages if it wasn't lagging so bad.


Mar 15, 2007
6:26 PM  
springnet wrote:

I'm up to 100 followers after just 3 or 4 days. Started using it on day 1 of sxsw.


Mar 15, 2007
6:28 PM  
mat wrote:

Very nice Andy. I love it when you run stats like these


Mar 15, 2007
7:06 PM  
/michael. wrote:

A Many Eyes treatment of the same data. Not as pretty. I sourced the data back to here and hope you don't mind.

http://services.alphaworks.ibm.com/manyeyes/view/SJjqGFsOtha61PEcjt5KF2-


Mar 15, 2007
7:46 PM  
Jeffrey W. Baker wrote:

Isn't this better thought of as a growth rate, i.e. first derivative of what's posted here?

http://tastic.brillig.org/~jwb/twitter.png


Mar 15, 2007
8:55 PM  
Alan wrote:

Not posting to be the internet jackass, but I think your data collection methodology is flawed, and this might explain the abnormalities you're seeing.

It's a mistake to assume that each sequential ID represents a single post in the system (unless you're privy to back-end details, in which case I'll shut up).

It's probably safe to assume that the number is an auto-generated sequential ID of some kind, but we don't know for sure by what amount this value is bumped up every time there's a post made to the system. An id of 8,281,991 doesn't mean 8,281,991 posts are in the system

So, here's one plausible scenario that's probably not "the truth", but it illustrates my point. Twitter starts as a small application with a single MySQL database server. The table that stores the messages has an auto increment primary key that gets bumped up by one for each row inserted.

At some point, one database server isn't enough. Twitter is an application that's going to be sensitive to slave lag, so the team decides to go with MySQL 5's multiple master feature which lets you have more than one master server.

One of the problems with multiple master servers is what happens if both servers receive an insert at exactly the same time that generates the same auto_increment primary key. The master/master synching voodoo pukes and corruption happens.

To get around this, the primary keys use a modulus. In a two master setup, one server generates odd keys, the other generates even keys. In a three master setup, each id is incremented by 3 with server 1 starting at 1, server 2 starting at 2 and server 3 starting at 3. Since the servers will always generate different primary keys, corruption is avoided.

With this setup, unless the traffic is evenly distributed among the database servers, the id is no longer an accurate count of how many posts are in the system at any one time.

Now, let's throw in another wrinkle. If you're smart, even if you only have a two or three master setup you use a higher modulus to make replicating new masters into the system a breeze. So, an application might only have two database servers, but their primary keys are being incremented by 5 each time, which will let the application scale up to five master servers without having to worry about re-jiggering the modulus each time you add a new server.

So, in this setup, even with even traffic distribution between the database servers, your "id as a representation of how many posts are in the system" is completely fucked. The November 23rd traffic spike could be when a system like this was bought online. The February 5 drop off could be a re-jiggering of a modulus that was, in retrospect, too high.

Other, less complicated, scenarios could involved a much high number being pushed into an auto_incrment field for some reason. Subsequent inserts would start at this new higher number.

I dig the graphs though (-:


Mar 15, 2007
9:18 PM  
Claus wrote:

Andy, the first surge is not on November 23. If you look closely, you'll see that use quadruled on November 21 and multipled by a factor of 7-8 the day after. November 21 was the day Twitter launched the "six word memoir" promotion.
Use remained flat for the rest of the year. The effect on March 10 is unmistakeable: 20 times the use of the day before.


Mar 15, 2007
9:41 PM  
tagami wrote:

I started on 11/18. I think the Thanksgiving phenom is because people are in motion during the holiday and are reaching out to their friends. That's what I did anyway...


Mar 15, 2007
10:53 PM  
Jeff wrote:

Hey, I joined 11/20. I'll take full credit.

Great graphs, thanks for posting them.


Mar 16, 2007
8:58 AM  
Kempton wrote:

Thanks for taking time to do the work to generate these graphs. And get the ball rolling to look inside Twitter a little bit.


Mar 16, 2007
10:02 AM  
Joshua schachter wrote:

Alan: They're almost certainly on a mysql backend with the default auto_increment stuff set. Which one should not do, for a variety of reasons:

http://joshua.schachter.org/2007/01/autoincrement.html

Joshua


Mar 16, 2007
12:13 PM  
rabble wrote:

First off, twitter does use autoincrement. They know they shouldn't, but they do. They also only have on backend database so far. So Andy's numbers look good from what i know of twitter's setup.

The question of why the jumps? Those dates correspond to when good IM support was added, API's where released and desktop apps where built, and when the signup process got streamlined. Before November you needed a mobile phone to join. Once that was dropped growth shot up.


Mar 16, 2007
2:13 PM  
Andy Baio wrote:

That sounds like a definitive answer to me. Thanks, Rabble.


Mar 16, 2007
2:33 PM  
Joshua Schachter wrote:

I assume they use something like PRIMARY KEY(msg_id) and KEY(user_id, twitter_dt).

under Innodb, if they had PRIMARY KEY(user_id, msg_id) and KEY(user_id, twitter_dt) it'd have to do far less disk seeks to do the join, because all of a user's messasges would be in contiguous pages.


Mar 17, 2007
4:07 PM  
Michal Migurski wrote:

Alan, if what you say is accurate, then there should be a sudden *and permanent* jump in the slope of the line, due to a more collision-proof use of the primary key space. Rabble the authority here, so I'm assuming that auto_increment is in use.

I'm especially interested in the sudden drop in growth over 2007, leading up to SXSW. Why is that? Will it fall back to that level once the Austin shine wears off?


Mar 17, 2007
6:53 PM  
Andy Baio wrote:

I've heard rumors that they were forced to block an entire country, which was abusing the SMS features and costing Twitter tons of money. That might explain the sudden dropoff.


Mar 18, 2007
7:11 PM  
Angulo wrote:

Thx a lot for the graphs. You do your research thoroughly.


Mar 19, 2007
9:24 AM  
Alan wrote:

Thanks for the twitter background info and the indexing lesson.


Apr 26, 2007
10:34 PM  
max wrote:

oh nice site dude,
i need some ideas like that!


Apr 29, 2007
1:09 AM  
Eli wrote:

It would be cool to see how this lines up with big stories about twitter, like John Edward's twitter account annoucement.


Jun 3, 2007
7:23 PM  
Tim wrote:

This is the first time I've heard of twitter. It's actually fairly addicting once you get in and start reading the posts.

Do you feel that the novelty will wear off after a few months (July / August)? Or, is this a mobile app that will help drive the development of more user generated content?

Thanks for the post...nice stats work. I'm curious to see how things go after the sxsw peak wears off. Long tail or not.


Jun 6, 2007
8:21 AM  
erik wrote:

"Human Giant using Twitter at the MTV Movie Awards"

Aziz also mentioned Twitter at the Sasquatch music festival where he hosted the main stage on the second day (sans the rest of the Human Giant people or MTV for that matter).


Jul 31, 2007
5:49 AM  
Create wrote:

Very informative data. Long tail at work indeed.


 

Leave a comment





Waxy Links
Ads via The Deck
May 22, 2008
Guitar Hero hacker plays Smells Like Teen Spirit without game console — I've seen several mods turning the guitar into an instrument, but never seen a guitar solo on one (via)
Philipp Lenssen on using hoax news to boost search rankings — recent story about a teen paying for prostitutes with dad's credit cards was a hoax to boost a finance site's pagerank
May 21, 2008
The New York Times' TimesMachine — back after an aborted launch that I missed out on, this is great archival work
Jamie Livingston's Photo of the Day — he shot a Polaroid every day until his death from cancer in 1997 (via)
Sneak peek at the new Last.fm redesign — too early to judge, but a glimpse where they're headed; announcement and more screenshots (via)
Slow-motion videos of people getting punched in the face — background on the clips, shot at 1000 frames per second (via)
Offline Oracle Ouija Board — summon the Internet spirits with a QWERTY layout and cursor control (via)
Huntington Hartford, kooky millionaire, dies at 97 — he blew $80 million of his inheritance on crazy schemes like a handwriting institute and automated parking garage
Mike Sacks' Photos of TV — "funk master friendly"
Funny Farm — old, but new to me; a dangerous time-waster, spoilers here
May 20, 2008
Amateur footage of the China quake moments after it ended — terrifying beyond belief
16% of US science teachers are young-Earth creationists — they believe humans were created in the last 10,000 years; won't someone think of the children!? (via)
Chuck Klosterman's list of the most fanatical fanbases — Tori Amos is at #2, which sounds right to me; nice gallery of uber-fans (via)
Errol Morris reveals the backstory behind one Abu Ghraib photograph — fake smiles, Weekend at Bernie's, and the coverup behind a prisoner's death during interrogation
Rory Root, owner of Berkeley's Comic Relief, dead at 50 — nice tributes from comic greats; the best comic book store I've ever seen, his passion was clear on the shelves
May 19, 2008
How to Use Nico Video — Japanese video site from 2chan's creator requires registration to view videos, so I'm slogging through it
YouTomb's tracking who ordered YouTube takedowns — recent takedowns are on the homepage of the MIT project (via)
Threadless Prints — limited-edition 18"x24" posters based on shirt designs (via)
Web Trigrams, visualizing three-word phrases from Google's corpus — gorgeous infoviz; more of his work from the massive dataset (via)
Homer Simpson says, "Beat it, Waxy!" — from the archives
SNL's Japanese version of "The Office" — "It's funny because it's racist."
Music Bounce — quirky game built on music patterns (via)
Vice-Presidential Prognosticatin' — Obama/Prime 2008
Drew Burrows' virtual girlfriend in bed, a lonely art installation — reminds me of Daki Makura hug pillows and bedsheets (via)
Webmonkey relaunches, again — this must be the third or fourth time the site's come back from the dead
May 17, 2008
The London Tower Bridge on Twitter — I hope someone wires up a Coke machine to Twitter next
Why We Twitter, academic paper from 2007 about microblogging — nice network analysis and node graphs towards the end; Scoble's tying disparate groups together
Wired News on the new Soulseek client for jailbroken iPhones — decent speeds downloading music and it imports into your iPhone music library when done
May 16, 2008
Bill O'Reilly's Producer — look at the other side of the camera
"Things Younger Than McCain" creator on the subject of ageism — unlike sex and race, your age affects your ability to lead

Andy Baio lives here. Some rights reserved, for your pleasure.