Thursday, May 29, 2014

Big Data vs. HPC

I wrote a blog post awhile back on "HDFS vs. Lustre".

The primary point of that post was that it was not reasonable to compare HDFS to Lustre.  Although I have never worked with other networked file systems like GPFS, Panasas, and pNFS, I believe the same argument can be applied to them as well.  Those networked file systems serve such completely different purposes and have completely different architectures that doing an apples to apples comparison is difficult if not impossible.

So I saw this article recently on Datanami, "Making Hadoop Relevant to HPC".

I felt the need to discuss many of the comments discussed in this article.

Lockwood argues, is that Hadoop “reinvents a lot of functionality that has existed in HPC for decades, and it does so very poorly.”

I can agree that Hadoop reinvents some functionality.  Most notably job scheduling and resource management is something HPC has done for a long time.  However, to my knowledge, HPC has not had a scheduler/resource manager that tightly integrated the filesystem with the job/task scheduling itself.  Therefore the need for the Hadoop community to make their own resource manager.  If you want to criticize the Hadoop community for not using the currently available open source resource managers and writing a plugin?  Ok, that's decently fair.
For example, he said a single Hadoop cluster could support only three concurrent jobs simultaneously. Beyond that, performance suffers.
I'm not really sure where the "three concurrent jobs" comes from.  This makes no sense to me.  I suppose it's possible that Hadoop's default scheduler elects to give priority to jobs differently than what is expected from a traditional HPC scheduler, but that's easily rectified through some mods to the priority queue algorithm.

I can believe that performance may suffer as you add more and more users.  After all, HDFS daemons sit on each node and may get busier and busier as you have more users.  However, I could make the same argument of traditional HPC file systems.  The more and more users you add to them, the busier the file system gets.  At the end of the day, you can only pump so much data through a network link.
Lockwood maintains that Hadoop does not support scalable network topologies like multidimensional meshes.
While technically true,  Big Data applications are programmed and designed in a completely different way.  They may not necessarily benefit from such advanced network topologies.  It's possible Lockwood has some specific applications he's thinking of that could benefit, but I would disagree with this statement for the general problems being handled.
Add to that, the Hadoop Distributed File System (HDFS) “is very slow and very obtuse” when compared with common HPC parallel file systems like Lustre and the General Parallel File System.
Now this comment I'm going to take a little more time to discuss.  Reiterating some of my points from my earlier "HDFS vs. Lustre" post, this is comparing apples to oranges.

The correct comparison is "MapReduce over HDFS/Local Disks vs. MapReduce over Lustre."  This is the real comparison.

MapReduce creates many small files during it's shuffle phase.  Does Lustre/GPFS perform well with small files compared to local disk

MapReduce performs many random-like seeks/reads during its shuffle phase.  Does Lustre/GPFS perform well with random reads compared to local disk?

When your data problem exceeds system memory and you need to spill contents to disk temporarily, will temporary scratch spills be faster to local disk or a networked file system?

I could go on and on and on with this argument.

The point is, is HDFS not as flexible as Lustre or GPFS?  Yes.  But does it serve its purpose better than Lustre/GPFS?  I think the answer is yes it does.

Hopefully in the near future I will be able to point to online published results illustrating this fact.

Sunday, March 9, 2014

Cardinals 2014 Roster Depth

The Cardinals just announced the signing of Aledmys Diaz to a contract.  I love this move.  In combination with the other moves the Cardinals made last season, this gives the Cardinals incredible depth.

The infielders the Cardinals took to the World Series last year were:

Allen Craig, Matt Adams, Matt Carpenter, Kolten Wong, Daniel Descalso, Pete Kozma, David Freese
Freese is gone, traded to the Angels.  Peralta is in.  Mark Ellis is in.  Presumably Diaz will be in soon.  Effectively, they replace Descalso and Kozma.  Regardless of who ends up playing second base full time, that's a much deeper bench.

The outfielders the Cardinals took to the World Series last year:

Matt Holliday, Jon Jay, Carlos Beltran, Shane Robinson

Carlos Beltran is gone, Peter Bourjos is in.  Once Oscar Taveras is called up, effectively Jon Jay will be put on the bench to replace Shane Robinson.  That is again, a much deeper and more talented bench.


Saturday, March 8, 2014

Another Great Cardinals Long Term Deal

The Cardinals continue to impress me with the long term extension deals they make on their players.

Today they announced a 6 year contract extension with Matt Carpenter for $52 million.  So that's $8.66 million a year to eat up Carpenter's three years of arbitration and two of his free agent years.

Matt Carpenter had an MVP caliber year in 2013, leading the National League in runs, hits, and doubles and placing fourth in MVP voting.  Even if he doesn't perform quite as well as he did in 2013, it's still a solid signing and they didn't stretch their dollars too much.  The Cardinals will get him through his age 33 season.

As a comparison, Dan Uggla got $62 million for 5 years by Atlanta in 2011 when he was 31.  Omar Infante signed a $30 million contract for four years and he's 32.

It follows up some other great signings, including Allen Craig for 5 years and $31 million.  That contract takes Craig into his age 32 season.  Allen Craig isn't going to surprise you as a superstar, but he's a solid offensive talent.  For just over $6 million a year, he's a far better value than what you can get on the open market for a first basemen/outfielder.

They signed up Yadier Molina for a five year $75 million extension in 2012.  That's basically $15 million a year for the best catcher in baseball through his age 34 season.

My favorite recent signing was the one for Adam Wainwright at 5 years for $97 million.  Given the huge $150+ million contracts lately given to Zack Greinke, Clayton Kershaw, Justin Verlander, and Felix Hernandez, it looks like a steal.

Naturally, no free agent contract signing is without risks.  However, the Cardinals regularly seem to play things smart.  They spend non-outrageous  sums of money on high reward/risk ratio players.  Some will not work out but they seem to work out more often than not.

Update 3/9/14:

And it gets even better, w/ the Cardinals signing Aledmys Diaz to a contract.  The Cardinals team depth is beginning to look crazy good.

The N=1 Problem

Recently read this article from ESPN about how the Angels are trying to rebuild their minor league system.

http://m.espn.go.com/mlb/story?storyId=10470778&src=desktop

One of the subtle reasons I love reading articles like this is that at the core, major league baseball teams are no different than other national or multi-national corporations.  All the same management, mentoring, training, recruiting, and retainment issues all organizations face are the same in baseball as everywhere.  It's just that when spoken about in a baseball context, the article is way more interesting than some droll tale of organizational synergy.
 
There's two chunks of the article I love the best:

Most of the lessons of the sabermetric revolution are based on what's called large-N analysis: looking at all the players who ever played and finding, in millions of data points, answers about player tendencies and optimal strategy, and meta-answers about the reliability of statistics. But developing a prospect is an N=1 problem: Each player's combination of skills, genes, experience, health, neurology, psychology, size and style makes him unlike any other player.
then later


How a coach teaches pitchers to back up a base isn't, ultimately, all that important. What's important is that no coach has to spend more than two minutes of his life thinking about it. That frees him to focus on the N=1 problems
In other words, if a coach has to waste his time dealing with "stupid stuff", then the coach can't concentrate on what's important, namely teaching the player what they need to be taught to reach the next level.

I can't help but think about this within the context of a lot of major companies.  Every employee will have different opinions on what are "annoyances" or "interruptions".  It's likely impossible to remove all of them for every employee, but the hope is that most organizations limit it to a N=2 or N=3 problem for most employees.  Unfortunately, I suspect many employees are dealing with N=9 or N=11 problems.



Thursday, February 13, 2014

HDFS vs Lustre

There's been discussion out there about comparing the HDFS filesystem to a traditional parallel filesystem like Lustre.  The problem is it's really difficult to compare apples to apples.

As an example, I saw a white paper awhile back (sorry, I can't find it online) that compared HDFS to Lustre.  HDFS beat Lustre in this person's performance tests by a good margin.  After digging into the paper I saw why.  This fellow ran Lustre over a 1 GigE ethernet network. 

Is this a fair test?  On the one hand it isn't because Lustre is a network based filesystem.  If you simple choose to bottleneck Lustre, of course it will lose.   On the other hand, it's a fair test, because it uses the same hardware most use with HDFS.

So lets say we replaced the GigE with Infiniband.  Would it now be a fair test?  Perhaps its slightly fairer, but HDFS people can say HDFS wasn't designed for more expensive hardware and therefore doesn't take advantage of it.  In the case of Infiniband, HDFS isn't using RDMA during replication.

I don't know the right comparison.  However, HDFS vs Lustre may not be the correct comparison to think about.  At the end of the day, I could probably concoct an HDFS setup that will always beat a Lustre setup and vice versa.

I believe thinking about this as HDFS vs Lustre isn't the right approach.  It's really Hadoop Cluster vs HPC Cluster.  At the end of the day, while Hadoop is famous for handling large data, the reality is because of shuffle/sorting/scheduling/etc. in Hadoop, it also reads/writes tons of small files.  The memory for a Hadoop Cluster vs HPC cluster may also be different.  That affects spilling of data, page cache, etc.

Update: See "Big Data vs HPC" follow up.
Update: See "HPC vs Big Data" follow up.

Update 6/2/15:

Not so long ago I was talking to someone about the HDFS vs Lustre comparison.

Many people have done HDFS vs "Some Networked Filesystem" experiments.

However, I think these experiments are inherently flawed.  The experiments always look something like this.

Datanodes
8 nodes
4 SATA disks
8 core
32G RAM

Networked Storage
4 nodes
8 SATA disks
8 core each
32G RAM

with additional hardware details beyond this.

The comparison will be HDFS using the Datanodes for data & map reduce.  Then it'll be a comparison to the Networked Storage, also using the Datanodes as the computation facility.

Do you see the inherent problem in the above comparison?

.
.

It's staring you right in the eyes.

.
.
.

It's an 8 node test vs a 12 node test.

It's a 256G RAM test vs a 384G RAM test.

This isn't to say that the comparison is poorly done.  But this is part of the inherent problem of comparing HDFS vs Networked File systems.  What is a fair comparison?



Wednesday, February 12, 2014

Big Data vs. HPC/Supercomputing

There's been a lot of articles about what is "Big Data" and how does it compare to traditional Supercomputing and High Performance Computing.  I thought about it, and devolved it into a simple mathematical statement.

In Supercomputing / HPC

Computation Time >> IO Time

and in Big Data

IO Time >> Computation Time

The architecture of the hardware, the networking solutions, the software you use, how you design your software, etc. etc. is centered around this simple statement.

Monday, January 20, 2014

High performance through strange correlations

Awhile back I learned of a really interesting statistic in Malcolm Gladwell's book Outliers.

In some standardized tests, test administrators like to ask students survey questions before the test so they can try and correlate test performance to other social factors.  These questions are pretty normal, things about your social life, friends, and family.  However, this survey is so long and tedious that many students just give up and don't even complete the survey before taking the test.

What several researchers found was that there is a strong correlation between percent of the survey completed and performance on the test.  Students that complete a higher percentage of the survey perform better on the test.  Note that it's not how they answered the questions on the survey, simply the fact that bothered to complete to the survey.

The suggestion is that patience, diligence, and work ethic are actually more important than smarts.  If you can stomach through this tedious survey, you've probably stomached through a lot of homework assignments to be able to learn the material well.

Not so long ago, I was on a committee to review a contract for a procurement.  While reviewing the bids, I noticed that there were bids that were horrifically bad.  At the same time I noticed the bad bids were much shorter in length than the better bids.  In fact, by the time the committee was done with the selection process, the ranked order of the bids was strongly correlated to the thickness of the printed bids.  My recollection is that if only two bids had been flipped in their ranking, the correlation between thickness of printed bids & rank would have been perfect.

It got me thinking, the correlation was similar to the correlation with the standardized tests.  The companies that put in way more effort and time into their bids, probably cared about their bid a lot more.  At the end of the day, the fact that they cared about the bid a lot more, probably meant they wanted to win the contract a lot more, and meant they would care about fulfilling the contract a lot more.