About a decade ago I tried to create an alternative to Google, using the Dewey Decimal System to group web domains by subject matter. My vision was that the website was a library, and domains were individual books. Rankings were a mixture of domain reputation and cultural significance.

The OCLC is very protective of their IP, and likes to sue.

Needless to say, my project was shut down before even getting off the ground.

I'm working on reviving the project, under a new classification system (The Open Categorization System), and would love any help I can get.

The project can be found here: https://github.com/ki4jgt/Open-Catalog-System

you are viewing a single comment's thread
view the rest of the comments
[–] 6 points 1 day ago (4 children)

I'm curious as to why you'd limit this to the domain level. A university for example might have thousands of URLs in it with wildly different subjects:

  • university.tld/math/student-name/thesis-on-mathy-subject/
  • university.tld/journalism/student-name/big-story-about-politics

How would your system account for this?

  • source
  • hideshow 4 child comments
  • [–] 2 points 1 day ago (3 children)

    I'm guessing that people would just find the content with the site's own navigation / search mechanism.

    I don't really see how it would make sense to replicate URLs below a domain level. Or how it would be manageable.

  • source
  • parent
  • hideshow 3 child comments
  • [–] 1 point 22 hours ago* (2 children)

    and I don't see any value in limiting classification to the domain level. How would one classify wikipedia.org in this scenario? Would it not make more sense to define an open standard that'd leverage this system but allow domain managers to define it themselves?

    To take my university.tld as the example, that university might host a file called at https://university.tld/ocs.json that looks something like this:

    {
      "/": "EDU"
      "/math": "MAT",
      "/journalism": "JOU",
    }
    

    (Heads up to OP: there's no journalism in your current spec. That feels like an oversight.)

    This would allow the high-content site admins to classify parts of the site differently. The /ocs.json file might even be dynamically generated for complex sites like Wikipedia.

  • source
  • parent
  • hideshow 2 child comments
  • [–] [S] 1 point 18 hours ago* (1 child)

    Wikipedia would be inf.ref.enc. Webster and Urban Dictionary? inf.ref.dic.

    Standard library locations.

    As far as the university goes, the system's incomplete.

    A library has books on several universities, but the book can only be in one place.

  • source
  • parent
  • hideshow 1 child comment
  • [–] 1 point 18 hours ago

    Well if you do decide to attach it exclusively to domains, you might want to consider developing a DNS TXT standard format so this sort of information can be stored in DNS and cached accordingly.

  • source
  • parent