Sunday, December 21, 2008

the three-card game


already on the 2-card machine, I played with the os installation. initially, I installed a 32-bit RHEL 5.2 from media, only to change my mind and go for the 64-bit fedora10, which should give me a slight performance edge.


I had a nasty day or two without the X windows/desktop, since the nvidia driver version 180.06 and then 180.16 did NOT install ok. whoever wrote that script seems not to have known a magic spell I finally found on nvidia forum, you know, something like:
"/usr/bin/nvidia-xconfig -a". a spell without which your computer only outputs 24 lines of text ;-). half a day, and a bunch of rpm's later I started feeling at home in my tcsh and gnome desktop. I installed the newest sdk for cuda without incident.

then the real fun began. the physical re-installation of the water-cooled cards.
my first leak... :-( well, all those of you who used the extremely short barbs in the BFG gtx280 H2OC kit, meant for SLI installation, know what I mean. the supplied short pieces of clear lastic vinyl tubing don't tightly fit around the barbs, the white plastic clamps aren't really working and, besides, the short barbs are a bit too wide and catch on the aluminum card backplate painted in black, when you try to mount them on the card. all those mechanical/hydraulic issues
can be solved by abandoning those too-short and too-wide barb fittings and using the 3/8 inch fittings, which are smoother and better quality anyway. just one thing: you have to shorten them so that the cards are close together. I put the cards in, measured the distance and cut off pieces of four 3/8" barbs with a carbon (diamond?) disk attached to an electric drill.
I smoothed the edges so they won't cut the hoses and everything started looking good again (although I'm missing the home depot's garden hose a lot! :-)




the LQ1000 case is very compact, like all mid-towers. for instance, I musty warn those of you who would like to use EVGA water-cooled 280s that this box DOES NOT WORK with them. they're too high (from PCIe bracket on mobo to the top of the card). you have to use BFG gtx280 H2OC cards. I was just lucky to have gotten them in my first iteration. they almost touch the big fan casing, so you have to route the cooling hoses over the low sectors of the cards. there is enough tubing in the Zalman kit for several waterblocks.



those hoses will likely touch the 24cm fan casing
which is ok. the small door covering the disks will touch their sata cabling. it's all barely doable. if you install the lowest/furthest x16 card, you'll always worry a lot about the sharply bent hose coming down toward the bottom plate of the case. but it will all work! I bent and even clamped the hose in a working system with my fingers and the flow rate diminished noticeably only when I used so much force that I thought I'll stop the flow completely. but it kept going and Zalman's flow rate alarm wasn't even triggered.


in this picture you can see that I removed the white plastic clamps between the graphics cards (cf. previous pix), and installed the home-depot clamps on the shortened 3/8" barbs. no leaks now.

the two-card game

Let's take a look at the first configuration I built, with two gtx280 gpus. I actually bought three cards, but got scared about the thermal limitations of my Zalman cooler (which are like those of Zalman XT external cooler): nominally only 500 W heat removed.
well, theoretically I had 3 x 236W of heat just from the 3 cards! so my system, on paper, was limited to 2 cards... but I figured that all the heat doesn't go into the coolant.












the de-gassing of the cooling system has to be done with all the waterblocks below the pump.
no leaks were found. if you use Zalman, remember not to give up (like someone on some forum did)
after a few beeps and disconnections of the pump. this is a normal behavior. before most air is gone you'll have to restart the system up to ten times. I recommend to turn the waterblocks as much possible during this; it helps the air escape.





I like this shot, it shows perhaps the first supercomputer built from pieces of nylon-reinforced garden hose from home depot :-) the wide spacing between the cards is due to the location of the 2 fastest PCIe x16 slots.

intro & specs

the system I'm going to describe is my first step on a new path into shared-memory parallel computing. in other words, here you won't find anything about how to improve your highscores and frame rates in Crysis 3.14 or Left 4 Dead 7.0. :-(

my Z-Machine (when I find a good name for it, I'll edit out ZMachine :-) was designed with these objectives in mind:

1. max power: max number of cards in one box,

2. well interconnected: cards sitting on pci express bus (x16 if possible) and not dependent on the relatively very slow gigabit ethernet switches (which aren't very broadband; in PCIe terminology they are x1 or x2 devices!)

3. quiet operation for office, not sever room setting

* * *

points 1 and 3 suggested water cooling, and when I started reading up on that subject, I was amazed that a little-known box called LQ1000 from a respected manufacturer Zalman has a nice cooling system integrated inside the box. although a bit expensive, it looks great (wine-colored gauges remind one of a bmw dashboard :-)

btw, the box looks like so

and not like this

prototype
from a 2007 trade show.

* * *

next: which cards? nvidia geForce gtx280 was my choice (240 cores!). what's interesting, water cooled cards by BFG and EVGA are factory overclocked. great!
I considered the newer versions of gtx260 with 216 cores but the price-performance calculus preferred gtx280. [I looked at the price and performance of the whole computer, not just one card!]

next: the motherboard and cpu. well.. that was kind of unimportant if my hopes as to the gpus were
right (-: so i settled on a run-of-the-mill quad-core intel processor...

* * *

it took me the last week of Nov 2008 to (over)design my machine while scanning the world for the following components (prices are approximate, in CAD):

ZMachine:

  • box and cooling: Zalman "ZMachine" LQ1000, with included cpu waterblock & whole cooling system in a midtower. $800

  • cards: 3 x BFG GeFOrce gtx 280 H2OC, factory-overclocked setting, 680 MHz main clock - $686 each at bestdirect.ca

  • motherboard: EVGA nForce 790i SLI FTW - $350 [FSB clock 1350 MHz, +15% overclocked PCIe.
    Good mobo, except for a tiny northbridge radiator fan, which becomes loud when nb is getting a workout by cuda applications. however, at the end of 2008 there simply were no better boards. I could (and maybe would) have opted for ASUS Striker II Extreme or a Gigabit board with i7 nehelem cpu (socket LSA1366), but then I would have a wrong Zalman cpu waterblock bracket, and the really insufficient PCIe throughput, about which later..]

  • CPU: intel quad-core at 2.83GHz (Q9550) - $300(?)

  • RAM: 2x2GB SLI-ready DDR3 1800 $ 440

  • PSU: Toughpower Thermaltake 1200W [has the required two modular +12V connectors to each of the three gtx280 cards, and is very quiet. pay close attention to the number of available power connectors if you construct a 3-way SLI!] - $430

  • 2 x 1TB Spinpoint harddisks from Samsung (quiet) - $240 (both) [I have a backup partition 250GB on the second drive, still don't know what I gain :-) since if the 1st disk crashes, the second is not automatically bootable... well I'll sort it out later]

  • 1 dvd-rom $29 [nice, quiet], kbd/mouse $10 ea.[spent too little? both aleady failing :-]

  • Samsung SyncMaster 2443BW, 24" 1920x1200 monitor. $350 [I like it, pivots around 2 axes, adjustable height. great contrast etc]

  • OS: Fedora 10 , x86_64, driver: nvidia 180.16 - $0 [I downloaded and installed 6-7 GB over the net without any physical media in one night]

  • CUDA v. 2.1 beta. [installs & works fine; I skipped compilation of those few examples that require some extra libraries]

GPU >> CPU

in early 2007, nVidia opened up the gates to a paradise. a free if not entirely open-source project called CUDA made its debut. it's a general-purpose graphics device computing, utilizing the massively parallel architecture of today's GPUs, or graphics processing units. in effect, all the recent nvidia cards became capable of carrying out parallel computational tasks. their raw power exceeds that of a CPU by a factor now typically ~10^2.

CUDA is best explained in this wiki page.
It is nicely illustrated using real-life applications in the
nVidia CUDA Zone. Typical speedups w.r.t. cpu are 5-100.

mythbusters were hired to illustrate the power of parallel processing and constructed a cute, monstual
parallel paint gun to paint mona lisa in less than 80 ms. it's fun to watch the monster and its 1024 paintballs flying slow-mo to meet their final destination: canvas.

without much exaggeration one can say that gpus could only be ignored so long - as soon as there's a solution that speeds up your program 100 times, you have no choice but to change to that track, no mater how comfortable your old one was.

for me, the old path was clusters and MPI. I started 10 yrs ago with a cluster of sun ultra5 workstations, and later continued with a cluster of custom built pc's in rack mounts.
hydra cluster, 5 GFLOPs
sample application of hydra
ANTARES cluster, 61-144 GFLOPs

MPI is a language (more precisely protocol and libraries that implement it, for exchange of data between nodes of a cluster). physical exchange was facilitated by commodity gigabit ethernet switches that became affordable about that time.


clusters were great, and essentially most today's supercomputers are built like that: farms of dozens to tens of thousands of machines hooked up by relatively slow interconnects. distributed memory and distributed processing power. which is ok for some problems, like hydrodynamics w/o radiation transfer or self-gravity in astrophysics, or frame-by-frame movie rendering and postproduction in a studio.

so clusters were great, but not trouble-free. you had to wait (hours or days, depending on how ambitious your computation was!) for the requested number of processors on some big national supercomputer to be allocated to your simulation. or you could decide to build your own little cluster, if you had money, place and time for it. that made more sense to many, and could cost your grant agency 'only' $20k or so, unless you really needed smp (shared memory machine, then you had to multiply the cost 5-10 times.)

you and your associates could only run a system of a few dozen nodes at best. beyond that magical number, frequent individual component breakdowns, software upgrades, and so on, needed to be taken care of by a professional sysadmin or technician (which you could not afford; so you used
your nodes praying they don't fail, and did not repair those that eventually did.)

on a small cluster, scientific long-term simulations could in practice be done in 2D but rarely in 3D, unless you were very lucky with your problem and/or very patient..

* * *

let's skip to 2007 then. why is a gpu hundred times more powerful than a cpu?
today, both are capable of parallel computation, since they have multi-core structure.
each gpu core is far less advanced on the control/vectorization side a bit slower.
but your ~$350 4-core intel processor is no match for 216-240 cores of a ~$350 nvidia gpu, on the newest G200 card series (nForce gtx260, gtx2800, and from january 2009 also a 2-gpu card gtx295 with 480 cores). ASSUMING YOU CAN harness the combined power of those gpu cores...

so here's the challange: to build and program a massively parallel system (with hundreds of cumputing nodes or cores) that is a bit more environmentally friendly than the old clusters: much less noise, much less total electrical power used, and finally much much more bang for the buck.
and that means: a supercomputer in a signle computer case, performing thousands of GFLOPs!
perhaps a thousand times the number of operations you could perform 10 yrs ago.

why cuda

why CUDA? the answer is simple: in my mother tongue (Polish) cuda means... "miracles".

why at all

I guess it's the envy...my teenage son built a little studio and I wanted to make some mess too













so I set up a little computer/hydraulic workshop in my library