• Status: Solved
  • Priority: Medium
  • Security: Public
  • Views: 307
  • Last Modified:

Algorithm to Score Global Quality-Quantity of Files (eBooks)

I would like to create an algorithm to score the "global quality/quantity" of eBooks of each subject.

Each subject has a main subfolder, like:

eBooks\Subject1\(...)
eBooks\Subject2\(...)
eBooks\Subject3\(...)
(...)

I would to create an algorithm/equation/formula to score them (subjects). I think it would be important such algorithm/equation/formula to rely on:

*) Total Size of each Subfolder (the higher the better)

*) Total Number of Files inside each Subfolder (the higher the better)

*) Total Number of Sub-Subfolders inside each Subfolder (the higher the better)

*) Maximum Folder Depth (the higher the better)

*) The Size of the Largest File Size (the higher the better)

*) The Size of the Largest File Size (the higher the better)


I tried many combinations but the resultant score were very absurd, except when:

Score = (Total Number of Files inside each Subfolder) * (Total Size of each Subfolder)

Do you know a better one?

Thanks.

Regards.
0
asgarcymed
Asked:
asgarcymed
  • 4
  • 2
1 Solution
 
ozoCommented:
It depends on how you want those factors to interact, but something like
total((size*depth)²) would seem to satisfy your criterion
Or you could might explicitly evaluate each of
*) Total Size of each Subfolder (the higher the better)

*) Total Number of Files inside each Subfolder (the higher the better)

*) Total Number of Sub-Subfolders inside each Subfolder (the higher the better)

*) Maximum Folder Depth (the higher the better)

*) The Size of the Largest File Size (the higher the better)

*) The Size of the Largest File Size (the higher the better)
(that looks like a duplicate)
and sum those individual scores, perhaps with some weighting factor

Given a few example folders and what you want their relative scores to be, we may be able to fit a function that orders them appropriately
0
 
asgarcymedAuthor Commented:
Please, download my CSV file (inside a ZIP) at:

http://tinyurl.com/ypv8t7

You will find a "Score" Column/Row, which corresponds to :

Score = (Total Number of Files inside each Subfolder) * (Total Size of each Subfolder)


You also will find that all numeric values are preceded with

«(zero or letter) »

because I do not know how to numerically sort the Columns/Rows inside a CSV file; and I do not want to sort it as alphabetic sorting:

1-10-100-1000-2-20-200-2000

instead of

1-2-10-20-100-200-1000-2000

If you know how to solve this; I also would appreciate your help ;)

Thanks.

Best regards.  
0
 
asgarcymedAuthor Commented:
PS - I do not why, but the "Experts Exchange" sometimes makes illegal characters; what can be awful in case of posting formulas/equations/algorithms/functions.

«(zero or letter) »
[illegal characters - why do they are generated????]
0
Never miss a deadline with monday.com

The revolutionary project management tool is here!   Plan visually with a single glance and make sure your projects get done.

 
asgarcymedAuthor Commented:
ozo - Do you have any news?

Thanks.

Regards.
0
 
JimFiveCommented:
I think what you would want to add is a "Weight" for each category to indicate how important each category is overall. So your formula would be something like (Weight1 * Category1) + (Weight2 * Category2) + ...

Also, By counting subfolders and depth separately it seems that you are giving extra weight to organization.

I would think just counting files or adding up file size would give you a quantity rating.  Beyond that you don't have any quality items at all anyway.  (Length of book <> Quality of book)

--
JimFive
0
 
asgarcymedAuthor Commented:
JimFive - Excellent idea! I was over-complicating! I agree with you 100%!!

Thank you very much for your suggestion!!

Best regards.
0
 
ozoCommented:
Didn't I say you could sum the individual scores with a weighting factor?
You never gave examples of scores from which we could determine what weights might work best to produce the desired order or whether interactions between categories would need to be taken into account.
0

Featured Post

Free Tool: IP Lookup

Get more info about an IP address or domain name, such as organization, abuse contacts and geolocation.

One of a set of tools we are providing to everyone as a way of saying thank you for being a part of the community.

  • 4
  • 2
Tackle projects and never again get stuck behind a technical roadblock.
Join Now