Showing posts with label shard. Show all posts
Showing posts with label shard. Show all posts

Monday, February 09, 2009

Speaking at Mysql Conf 2009: Architecture and Technology, Cloud Computing, LAMP, Replication and Scale-Out

I'll be going into detail what is Sharding, how to Shard, pitfalls of Sharding, performance/throughput gains, shard roles, and performance scaling in general. I hope to make this the most comprehensive talk to date on the subject in 45 min.

The topic is called Scaling a Widget Company. I'll detail how I setup the data layer for Rockyou. How many transactions per second Rockyou is at, what the infrastructure is comprised of, how 99.999% uptime is achieved and hopefully get into BCP which I probably will not have time to go over.

If you want me to focus on specific aspects on the subject of shard'ing let me know and I will :).

Thursday, August 07, 2008

Capacity Planning, Architecture, Scaling, Response time, Throughput

First of all let me start off saying that I learned a lot of Capacity Planning from two people. Jozo Dujmovic, and John Allspaw-who by the way is coming out with a book.

Capacity != Performance
. You may have the capacity to do a bubble sort but a bubble sort is still a bubble sort.

Really to Scale you need to know when your application will break. I have a tool set to help determine what application is producing what SQL and use that to figure out which SQL is producing the most load on the system. Some common tricks I do is put the execution path automatically as a SQL comment, then sample the FULL Processlist to build a graph on what application, function, SQL pattern is the top load.

On top of that I use Ganglia to trend the use of each mysql box. Key metrics that I use to determine capacity.

From iostat -x

I/O wait
atime
svctm

If the service average is trending towards 20% I/O wait I know that that is a hard-limit for my server configuration that will cause slave lag.

Jeremy Cole has a good write up and a tool for getting I/O stats that iostat itself does not expose.

If the atime (Response time) is growing, I know that the overall SAN LUNS are saturated. On top of that SAN LUNS typically have larger svctm cutting overall throughput with how MYSQL works. On a side NOTE I despise using SANs for mySQL. Why well that's another post.


Then I have ganglia configured to monitor everything for SHOW GLOBAL STATUS, but really I only look at the following

Com_delete
Com_insert
Com_replace
Com_select
Com_update
Key_reads
Questions
Connections
Threads_created
Slow_queries
Handler_read_next
Handler_read_rnd
Handler_read_rnd_next
Handler_rollback
Innodb_buffer_pool_read_requests
Innodb_buffer_pool_reads
Innodb_pages_created
Innodb_data_pending_reads
Innodb_buffer_pool_read_ahead_rnd
Innodb_data_read
Innodb_data_written
Innodb_row_lock_time_avg
Innodb_row_lock_time
Table_locks_immediate
Table_locks_waited
Sort_merge_passes
Sort_range
Sort_rows
Sort_scan


Next I take the techniques I learned from John Allspaw and build a 3rd order Polynomial and verify that my R^2 is in the 98%tile to see when I need to add more servers. So far so good. Now I have a rough idea when I need to add more servers-a capacity plan. (The techniques involved are various ratios of Users per Application per Shard, busy time, more junk like that)

Now your Architecture allows you to Scale, by ensuring a High Throughput at a low Response Time.

I personally use a architecture that I've started on since 1999-Shard'ing. Brad Fitzpatrick when building Live-Journal really made this concept popular.

With my Federation strategy I've been able to scale some of the most toughest dynamic applications linearly by just arbitrary adding more servers. It takes 5 min. to deploy new DB servers.

So, in summary to capacity plan you need to know how the system works, monitor it and trend it. To scale: your database architecture needs to meet the needs of the app. Are you read or write heavy or both? Do you have a lot of concurrency? Does your app do a lot of sorts? Does your app do a lot of ranges? Is it all of the above? Design to meet the needs, benchmark, know when it will break, and have a plan to recover before it does.