asked on

Parallel Computing: Clusters or Graphics?

I'm working on my AI research, which could benefit greatly from parallel computing. As I'm sure there's someone here who knows more than I do on the different platforms, which one would bring better results? (I know I could use both... but time's not really on my side.)

- nVidia CUDA [w/ GF8800s]
- OpenMPI [w/ about 50-100 computers, running P4D]

Some pages linking to benchmarks would be useful.

[Sorry for the low point count, I ran out. :(]

Thanks!

Callandor

One Tesla card supposedly can run at 518 gigaflops http://xtreview.com/addcomment-id-2756-view-Nvidia-Tesla-c870,D870-and-s870.html+tesla+nvidia+benchmarks&hl=en&ct=clnk&cd=2&gl=us, which is compared to the throughput of 40 x86 processors. There is a 4-card version for servers that is that much more powerful. Graphics cards are designed for parallel processing of textures and have a much higher transistor count than cpus, so it is not surprising that they can outperform general purpose processors for certain applications.

holobyted

ASKER

What would the higher-end Tesla card compare to? Ie, one "normal" Tesla card compares to 40 x86 CPUs (which CPUs?), what would the other be?

ASKER CERTIFIED SOLUTION

Callandor

membership

This solution is only available to members.

To access this solution, you must be a member of Experts Exchange.

Start Free Trial

holobyted

ASKER

How would 35 Pentium 4 D @ 2.00GHz compare? What would be the "rated" Xflops? Assuming peak performance.

Callandor

Are you talking about a real processor? I don't think there was a Pentium-D that ran at 2GHz. If you want to compare somewhat current cpus, use this: http://www.tomshardware.com/charts/cpu-charts-2007/pcmark-2005-cpu,382.html?p=1272%2C1271%2C1307%2C1306%2C1266%2C1274%2C1247%2C1242%2C1240%2C1302%2C1233%2C1300%2C1304%2C1254%2C1229%2C1297%2C1293%2C1296%2C1291%2C1290%2C1289%2C1279%2C1288%2C1312%2C1218%2C1309

holobyted

ASKER

If I recall correctly, P4D's went up to 3.2GHz... According to Wikipedia though, (http://en.wikipedia.org/wiki/Pentium_D), you're right.

What would be the approx. flops be for such a cluster? I'll try getting in touch w/ the owner of the 35 CPUs so I can get a real speed value. (Running OpenMPI)

Callandor

A single PentiumD 3.2 clocks in at about 600 megaflops, so 35 of them will be around 21 gigaflops. The PentiumD cpus are much lower in performance than the newer Core2 cpus, easily trounced by even AMD's X2 offerings.

holobyted

ASKER

Wow. That's actually pretty depressing... 35 systems can't even match up to one graphics card. Too bad CUDA is a pain to implement...

Callandor

Modern graphics cards are very powerful, and the ability to use them in non-graphics applications is very nice. Think about a $200 card giving you the power of 10 modern cpus - that's quite a good deal.

holobyted

ASKER

Yeah, I know. What's the GFlops on a "normal" GF8800 though? The tesla is outstanding, but that's cause it's a "small supercomputer for your workstation."

Callandor

It's about the same - 500 gigaflops: http://en.wikipedia.org/wiki/GeForce_8_Series#8800_GT, though I don't know if all of that is available for number crunching.