[2 days left] What’s wrong with your cloud strategy? Learn why multicloud solutions matter with Nimble Storage.Register Now

x
?
Solved

NoSQL or RDBMS? Web crawler

Posted on 2012-03-27
3
Medium Priority
?
1,309 Views
Last Modified: 2016-02-10
I need to crawl selected websites for about 5-10 different attributes.

For example lets say its a car website and on each page of the site there is information about a particular car for sale and it includes the vehicle make, model, year, price and etc.
I need all this information to be collected and stored in a database but since a good car sale website could have thousands of pages it can become a lot of data to collect.

I don't expect to have more than a few hundred words collected from each page so i think it would be under 1KB of data per record i store.

At the moment I don't know if i should be using NoSQL or a MySQL database since i will have an insane amount of rows/records created.

Any thoughts on going one way or the other? I need to do certain data manipulation on all the rows/records such as organizing the car by price from highest to lowest and etc.
0
Comment
Question by:checkmofoshoduno
[X]
Welcome to Experts Exchange

Add your voice to the tech community where 5M+ people just like you are talking about what matters.

  • Help others & share knowledge
  • Earn cash & points
  • Learn & ask questions
3 Comments
 
LVL 24

Accepted Solution

by:
johanntagle earned 2000 total points
ID: 37774846
Either MySQL or MongoDB (that's the only NoSQL db I've used so far) should be able to handle your needs, though MongoDB provides better horizontal scaling via sharding, should your data be really that huge.  If everything can be contained on one server, I would think the decision would depend on whether or not you can pre-define all the fields you need.  If so, I would personally use MySQL because I find querying via SQL more straightforward vs having to deal with mapreduce and the like.  But if you need to dynamically store different field names, then a NoSQL database is the way to go.
0

Featured Post

Moving data to the cloud? Find out if you’re ready

Before moving to the cloud, it is important to carefully define your db needs, plan for the migration & understand prod. environment. This wp explains how to define what you need from a cloud provider, plan for the migration & what putting a cloud solution into practice entails.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

What do responsible coders do? They don't take detrimental shortcuts. They do take reasonable security precautions, create important automation, implement sufficient logging, fix things they break, and care about users.
Backups and Disaster RecoveryIn this post, we’ll look at strategies for backups and disaster recovery.
Starting up a Project
In this video, Percona Solution Engineer Dimitri Vanoverbeke discusses why you want to use at least three nodes in a database cluster. To discuss how Percona Consulting can help with your design and architecture needs for your database and infras…

649 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question