Go Premium for a chance to win a PS4. Enter to Win

x
?
Solved

Questions about designing a data mining website and crawler

Posted on 2008-10-15
2
Medium Priority
?
782 Views
Last Modified: 2013-12-09
I have a few questions about a project I am working on. Being fairly new to the whole idea I decided to read up and found a great deal of information. The project has been designed and I have started writing the code for it but there are some issues that keep coming up.

Firstly, the bulk of of code comes in the form of class libraries that contain AI, rule processing and inference, database access, compression, etc. The crawler is also a class library that will reference the other libraries. The crawler will most likely be initialized by a console application or winform so that it will run outside of the asp.net session (any thoughts on running it from the asp.net website?).

So the first question is:
How can I control, manage, communicate with the web crawler when its running without using remoting or tcp client/server? Would I have to use a web service?

Second question is:
Is there a better approach to this design?

As it stands now I would like to have the crawler sit waiting for jobs to come in and then store the information into the database. I do not want the website to have to reference the libraries but still be able to access the data from the crawler and manage it as well.

My main concern is that if I use the scheduler I wrote to schedule the jobs and start the crawler the crawler will close out when the session from the site has ended. I am sort of lost on how to continue with this part.

I appreciate any help I can get and if I am being too vague just let me know and I will try to explain it in more detail and/or provide code snippets. Just as a side note, I am running SQL Server 2008, Windows Server 208 (IIS 7) and .NET 3.5 (Using Visual Studio 2008 to write it)

Thanks
Joe Wood
0
Comment
Question by:JoeDW
2 Comments
 
LVL 5

Accepted Solution

by:
wickedpassion earned 2000 total points
ID: 22739193
0
 
LVL 1

Author Comment

by:JoeDW
ID: 22749598
Wow, I really like the first link. It had tons of good information and I am sure it will keep me busy for a while. Thanks!!
0

Featured Post

Ask an Anonymous Question!

Don't feel intimidated by what you don't know. Ask your question anonymously. It's easy! Learn more and upgrade.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Originally, this post was published on Monitis Blog, you can check it here . It goes without saying that technology has transformed society and the very nature of how we live, work, and communicate in ways that would’ve been incomprehensible 5 ye…
Without even knowing it, most of us are using web applications on a daily basis.  In fact, Gmail and Yahoo email, Twitter, Facebook, and eBay are used by most of us daily—and they are web applications. We generally confuse these web applications to…
Explain concepts important to validation of email addresses with regular expressions. Applies to most languages/tools that uses regular expressions. Consider email address RFCs: Look at HTML5 form input element (with type=email) regex pattern: T…
The viewer will learn how to create and use a small PHP class to apply a watermark to an image. This video shows the viewer the setup for the PHP watermark as well as important coding language. Continue to Part 2 to learn the core code used in creat…
Suggested Courses

885 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question