Friday, August 28, 2015

Looking back at "The Anatomy of a Large-Scale Hypertextual Web Search Engine"

I recently read the famous paper "The Anatomy of a Large-Scale Hypertextual Web Search Engine".  It's a paper written by the Google co-founders Larry Page and Sergey Brin circa 1997/1998 about their web search engine research while they were students at Stanford.  The very first sentence of the paper summarizes its contents quite well, "In this paper, we present Google, a prototype of a large-scale search engine ...".

The paper is very interesting looking back on it 17-18 years after it was published.  I thought I'd comment on some of the fun things I read.

Improved Search Quality

November 1997, only one of the top four commercial search engines finds itself (returns its own search page in response to its name in the top ten results)
If the above is true, it is truly comical by today's standards of web search quality.

Major Data Structures

Throughout this section, Brin & Page continually do "bit stuffing" to save storage space.  Typically only done by those dealing with firmware, I find it a little ironic that they had to go to such lengths.  Given the amount of data they had to deal and the amount of hardware resources they had, it was obviously justified.  But it's sort of funny to think about it given today's data sizes and hardware resources that Google, Facebook, Yahoo, Bing, etc. have.

Servers to Crawl the Web

The original Google used a single URL server to serve lists to 3 web crawlers.  Insanely tiny by today's standards.  Of course, it was a much tinier web in the 1990s.

Social Consequences to Web Crawling

Perhaps the best part of the paper, Brin & Page talk of the social consequences of their crawler.  Most notably, some website owners were confused at what a web crawler was and why they were looking at their page.  Some would e-mail them asking questions ... some even called them.

Storage Requirements

Apparently the original Google had a compressed repository of just 53GB of data.  Insanely puny by today's standards.

System Performance

In addition, it took only 9 days to download all of the data on the web at the time.  It's not clear how many machines were at their disposal, but it did not appear to be more than maybe a dozen (as said above, they only used 3 for web crawling, and they note they used 4 for sorting the index).

"Advertising and Mixed Motives"

In this appendix section Brin & Page talk about the conflict of interest that search engines have when advertising is involved.  They specifically site the search of "cellular phone" as a keyword and say

It is clear that a search engine which was taking money for showing cellular phone ads would have difficulty justifying the page that our system returned to its paying advertisers. For this type of reason and historical experience with other media [Bagdikian 83], we expect that advertising funded search engines will be inherently biased towards the advertisers and away from the needs of the consumers.
It's ironic of course, b/c this is nearly the exact opposite of modern day Google.  A search for "cellular phone" on the site returned for me (in order)

  • An iPhone ad on apple.com
  • An ad for cell phones off a retailer site
  • An ad for Sprint
  • A Google Maps result for several retailers that sell cell phones
  • The Wikipedia article for "Mobile Phone"
This doesn't count all of the ads that are on the right hand column.

Sunday, August 23, 2015

Dinner @ La Folie in San Francisco, CA

I finally got the chance to hit up La Folie in San Francisco.  La Folie is one of San Francisco's most famous French restaurants, having been there since 1988.  It's been awarded a Michelin star since the first guide in San Francisco in 2007 and has a four star review from the San Francisco Chronicle.

La Folie offers either a tasting menu that's 5 courses (plus the amuse, petites, palette cleansers, etc.) or a "pick-n-choose" price fixe menu.  You can pick anywhere from 3 to 5 courses out of a list of about 20.  It includes appetizers, entrees, and desserts.  The only rule is you can't pick more than one entree item.

We elected to go with the pick-n-choose and we both decided to stay conservative and order just 3 courses, an appetizer, entree, and dessert.  That was a wise choice.  We would have been stuffed with more.  If you go with more than 3 courses, I'd recommended choosing dishes that are a bit more light and small (such as soup).  We did however, add a small course of a half of an ounce of caviar, just for fun.  So it was pushing 4 courses anyways.

So first we got some amuse dishes.

1) foie gras soup / cappucino


The waitress said this was one of the chef's specialties, a foie gras cappucino, although I think it I would liken it more as a foie gras soup.  The waitress even called it a foie gras soup to the table next to us.  I thought it was really tasty.  Admittedly, I haven't had enough foie gras in my life to discern the foie gras flavor, but it was a very tasty broth.

2) quail egg w/ corn soup


Next was another amuse dish, a small poached quail egg in a corn soup.  I thought it was an interesting dish overall.

3) russian osetra caviar w/ lobster potato blinis and creme fraiche


Next came our caviar tasting.  I was actually a bit surprised how much caviar there was, as I thought a half ounce was gonna be a lot smaller.  The picture perhaps doesn't do it justice.  I didn't know what lobster potato blinis was, but it apparently was just some lobster between two small potato-like pancakes.  It was an actual nice chunk of lobster claw, not just little lobster pieces which I thought it would be, so that was nice.  I thought the caviar was delicious. I probably would have enjoyed it more just by itself, although everything else was nice as well. 

4) day boat scallop, with uni, crème beurre noisette, braised young leeks - w/ corn


My date and I couldn't figure out what to have as an appetizer and we both really wanted the scallop, so we got the same one.  It was delicious.  The picture doesn't do this scallop justice as it was quite large and sizable.  Not the normal "entree" like scallops you normally see.  I had this feeling that corn was in season, as there were numerous corn dishes on the menu.

5A) Liberty Farm Duck Breast, Tokyo Turnips, Confit Cockscomb, Stone Fruit - Duck Jus - w/ Plums, Bok Coy, and Foie Gras


I got the duck breast as my entree.  It was cooked rarer than I think I've ever had duck before.  It was quite soft and far better than the duck I recall having a Cotogna.  There was a small bit of foie gras that was added (it's the small roll at the bottom of the photo) that was infused with some citrus flavors.  It and the grilled plums were particularly delicious.  I wasn't particularly enthused with the cockscomb, which had a bit of a muschroomy like texture to them (which I hate mushrooms).  I didn't even know what the ingredient was and had to ask the waiter at some point.

As an aside, I almost ordered the lamb as my entree.  I'm glad I didn't when we saw another couple order it at another table.  The rack was probably twice the size of the duck breast.

5B) Trio of Devil’s Gulch Ranch Rabbit, Chanterelle Mushrooms, Summer Vegetables, Natural Jus - w/ rabbit liver


My date got the trio of rabbit for her entree, which ended up with a bonus of an extra part of the rabbit.  Included were the rack of rabbit, a wrap of the rabbit loin, rabbit liver, and another part of the rabbit that I can't recall.  I thought the coolest part of the dish was the garden of vegetables lined up at the back of the dish. 

It was a lot more food than my duck breast.  The waiter actually told my date, "It's a lot of food, you don't have to try and finish it all".  I actually had to help her eat a nice chunk of this plate and we still couldn't finish.

6) summer berry soda


Then we got a light refreshing summer berry soda as a palette cleanser.

7A) Tasting of Melon, Watermelon Consommé, Cucumber Granite, Espelette - w/ Cantaloupe Sorbet


For my dessert, I specifically wanted something light.  Normally this dish comes with pineapple, but since I'm mildly allergic I asked them to remove it.  Overall, a nice refreshing sorbet with fruit.  It was a nice end to dinner.

7B) Mousse Au Chocolat, Smoked Chocolate Cremeux, Chicory Ice Cream, Cocoa paper


My date went and got something a bit heavier for dessert.  From what I tasted, also quite tasty.  The cocoa paper was particularly nice.

8) petite fours

And finally we got some petit fours.  The strawberry gelatin was particularly good.

The meal took about 2 hours total, which was perhaps a tad fast by most fancy restaurant standards.  It's perhaps part of the reason we were a tad stuffed after our entree course.   The heavier meat entree dishes definitely stuffed us.  The portions for every course at La Folie (with the exception of the dessert) were quite sizable, which was a surprise.  So for those looking to head over there, be careful of ordering more than 3 courses.


Friday, July 17, 2015

The All Star Game Should Be an Exhibition

I was talking with a co-worker the other day who told me that viewership for the MLB All Star game is down nowadays.  I used to love watching the All Star game, but my interest over the years as waned.  I believe it's for several reasons:

A) The rosters have exploded, so being an All Star doesn't mean as much as it used to.  Including those players who were injured, both the NL and AL rosters had 38 players on them each, well in excess of the normal 25 man rosters.  As a completely arbitrary comparison, the 2000 rosters had 35 and 34 players on the NL and AL rosters respectively.  In 1990 they were 29 for both the NL and AL rosters.

B) The game "meaning" something, means that players can't be loose and have as much fun as they used to.

Some of the best moments I can recall in All Star history are just fun moments.  The best players in baseball coming together for one day to show off their talents.

Two of the most memory moments I recall in the All Star game involved Randy Johnson.  One in which he psyched out John Kruk and the other in which he threw behind the back of Larry Walker.







Another great scene was when Torii Hunter stole a home run from Barry Bonds and Bonds playfully tackled him in center field afterwards.





The 2015 game had several fun highlights such as Jacob Degrom and Aroldis Champman striking out the side in their innings.  But for some reason, it just doesn't seem to be quite the same nowadays compared to years past.

Saturday, July 11, 2015

HPC Clusters vs Big Data Clusters: Two Different Worlds

Recently, I was thinking about why it's so hard for "HPC" cluster users to understand why "Big Data" cluster users do what they do, and vice versa. 

I wrote down this chart with a comparison of the software sometimes/often used on each:


Software HPC Clustering Big Data Clustering
Schedulers/Resource Managers Moab, Slurm, LSF, Torque, PBS YARN, Mesos
File Systems Lustre, GPFS, pNFS, PVFS HDFS
API "Framework" MPI, OpenMP MapReduce
Main Programming Languages C/C++, Fortran Java, Scala
Interconnect Infiniband, Myrinet, ... GigE
Higher Level Scripting ??? Pig, Hive


I could probably go on, but hopefully you get the gist of things.

Basically, everything listed under the "HPC Clustering" column isn't used on the "Big Data Clustering" column, and vice versa.

Here in lies the issue why the users of both don't understand each other.

I believe HPC cluster users look at the list on the right and immediately think things like:

  • "Why would you use HDFS, it's not a Posix file system."
  • "Why would you use Java, it's so slow."
  • "Why use GigE, that's so slow."
  • "Why did you write a whole new scheduler, why not use the schedulers HPC users developed years ago."
 In contrast, Big Data cluster users think nearly the opposite:

  • "Why would you use a Posix file system, that API/interface is ancient."
  • "Why would you use a networked file system, it's so slow."
  • "Why waste money on Infiniband, it's completely unnecessary to spend money on unused bandwidth."
  • "Why use MPI, the API is so complex, you can't develop programs quickly."
  • "Why use C or C++, the programming language is so complex, you can't develop programs quickly."

The problems users face are so different, that neither side can really understand why the other user would even bother to use the software/hardware that they are actually using.


So what happens when users in one world want to run in the other world?  I think what often happens is you hear "Can you port your code/application to work here?"  The answer is likely "No, that's not reasonable."  I believe you get these answers because most don't understand the difference between these two worlds because they don't understand the chart above.

So for those who are trying to mix environments, I think the most important thing to do is to try and accept the differences listed above and work for solutions that bridge the two worlds.

To some extent, that is part of my goal when developing Magpie (github).  Accept that the traditional HPC world isn't going to change and the Big Data world will not change either.  Better to try and get the Big Data world into HPC clusters with as little change as possible.

Update: See "Big Data vs HPC" follow up.

Tuesday, June 30, 2015

Final Fantasy 7 Remake: A less stereotyped change for Barret?

At E3 this year the long awaited high-definition remake for Final Fantasy 7 was announced.




I didn't think much about it, until I saw an article from Kotaku (here) that the game wouldn't be just an HD remake and it would have some changes.  Then another article (here) touching on how some scenes and components may be hard to pull off in HD without the "cartoony" look of the original.  The cross-dressing scene and slap-fight scene are two that come to mind that may be hard to do in HD.

It then occurred to me there was one component of the game that might get changed a lot.

Final Fantasy 7's Barret has been charged by some as racist (see wikipedia & article from kotaku).  It's been at least 15 years since I played the game, so I cannot remember everything about it.  I do recall Barret as the only character that cursed.  I also remember at times he used broken English/street slang while the other characters didn't.  Why?  I guess that's why they're called stereotypes.

Obviously, the game is now almost 20 years old and the world has changed a lot, and that includes in Final Fantasy games.  Final Fantasy XII's Fran (if she counts as Black, she is technically part of a half-rabbit race of people in the game) and Final Fantasy XIII's Sazh were quite normal.

So I'm wondering if a less "Mr. T" like character will be portrayed in the game.  We'll see.


Sunday, June 21, 2015

Employee Retention: The N = large vs N = 1 Problem

During a recent conversation a great analogy popped into my head on employee retention at companies that I thought I'd share.  It of course relates to my love of baseball. 

The Baseball N = Large Problem

The N = Large problem is quite basic.  Take data from a very large sample and use that data to determine some type of result (such as success/failure) with a new case not from the sample.  For those familiar with machine learning, this should sound very familiar.

The N = Large problem is widely used in professional sports.  In baseball, it is commonly used with the drafting/signing of amateur talent.  After analyzing the height, weight, speed, velocity of fastball, accuracy of throws, speed of swing, walk rate, strikeout rate, etc. of a player, teams can determine lots of things:


  • the probability of a player making it in the major leagues
  • the probability they will injure their arm throwing
  • the probability their performance will justify a $5 million dollar signing bonus
  • ... a $2 million dollar signing bonus
  • ... a $1 million dollar signing bonus

Teams are able to do this based on the large amount of historical data on players that have played in the league before and all the players that have played in the minor leagues and college.

The Baseball N = 1 Problem

In professional sports, especially baseball, each individual player's ability to succeed in the sport is never 100%.  The vast majority simply cannot jump from being an amateur to pro [1].  While the historical data can give you good guidance for the probability a player can succeed, individual coaching, mental training, physical training, and mentoring ultimately decide the level of success a player has.  Even for the most talented individuals, this can be the difference between a player making it to the big leagues or not.  This training is unique to each specific player.  Perhaps a player's swing is bad, or they need to get stronger to hit for more power, or their footwork needs to be better, or they need to learn how to throw a changeup, etc. 

As an example, in an article from ESPN they highlight minor league player Alex Yarbrough.  He is a reasonably talented player coming out of college.  However his throwing mechanics are so bad, he will never make it to the major leagues as a second basemen.  So the Angels must work on those mechanics to turn him into a major league capable second basemen.

This is the N = 1 problem.  Almost every amateur player has weaknesses that prevent them from playing at the major league level.  Each player's weaknesses are unique to them.  The work needed to fix them is unique to each player.  It requires special attention from coaches and trainers to deal with these weaknesses and help each player improve.

Employee Retention: The N = large vs N = 1 problem

In my opinion, there are two employee retention problems.  However, only one is often discussed.

The first retention problem is the N = large problem, retention programs to help all employees.  This is the most common type of employee retention issue that is discussed and the one most companies try to improve on.  It's easier to do and effects the most employees.  You can gather data about employees, conduct surveys, and determine best solutions for the employee population.  In all honesty, they are probably also the best bang for the buck retention ideas.  What are N = large retention solutions?

  • Lets give the employees a recreation room
  • Lets make ice cream parties for the employees
  • Lets add additional educational opportunities
  • Lets do an employee hackathon
  • Lets try to ...

etc.

However, there is an additional N = 1 problem to retention.  What is it?  It's concentrating on the needs that each individual employee has.  Every employee's pay, career goals, personal value in the work their doing, personal frustrations, etc. are all different.  And that must be managed employee by employee.  It's unlikely the efforts to solve retention in the N = large problem will apply here.

IMO, this is the retention problem that is typically not discussed or discussed very little.  After all, it's hard.  It requires managers and mentors (i.e. coaches) to look at each employee individually and determine what would be best to do for them.  That's really hard.  However, it may be the problem that needs to be discussed far more often.



[1] - According to Wikipedia, the last three players to do this were Mike Leake in 2010, Xavier Nady in 2000, and Jim Abott in 1989.  It's very rare.  Only 3 players in the last 26 years out of about 1000 players drafted a year.

Sunday, June 14, 2015

The Best Interview Answer I Ever Heard

There's been much written online about the best way to interview, the best candidates to look for, the qualities of top engineering talent, etc.

There is one singular quality I look for in any candidate, regardless if they are a system administrator, software engineer, or any technical position.  It's the ability to research and learn.

I was interviewing a system administration candidate for another group and the classic question I ask is, "When moving from administrating a few machines to 1000s of machines, what difficulties do you imagine you'll come upon?"

This particular candidate had actually setup and administered a small 16ish node cluster before and said something along the following:

When I setup this cluster, I realized FOO was running really slow.  I went online to see how other people solved the problem.  I found someone else who used pdsh to make FOO run better.  So I downloaded pdsh, set it up and FOO was working better.

I can't even remember what FOO was, but it was really irrelevant.  The candidate:

A) Realized something was running poorly or sub-optimal

B) Researched online a mechanism that would be better

C) Set it up/implemented the solution

D) The solution was deemed much better than the prior situation

Because my group had developed pdsh that was sort of bonus points for the candidate, but I absolutely loved this candidate's answer.  It was so simple and basic, yet illustrates exactly the quality that you want in an engineer.  Going online to research, learn, figure new things out, and find better solutions.  It's actually a quality that is often difficult to find in many candidates.