Based on a report by @LeChasseur and @armijnhemel
There is a large number of URLs that reference licenses and we have a good number of them as detection rules.
These rules are qualified as "is_license_reference"
- It would be better to mark them as "is_license_url" so we can distinguish them
- since there is a very large number of these possible we could optimize a secondary index and matching technique just for these
- some of these may match patterns such things based https://github.com/remy/mit-license or license badges
- for these we may have a special way to handle patterns
- some of these contain extra information at the referenced URL page.
- we do not want to live fetch them in scancode-toolkit, but tagging these with "is_license_url" means we could later have an optional step in scancode.io pipelines that could follow the referenced URL, fetch, save and scan them, including collecting copyrights and other useful information that may live there.
Based on a report by @LeChasseur and @armijnhemel
There is a large number of URLs that reference licenses and we have a good number of them as detection rules.
These rules are qualified as "is_license_reference"