The discoveredpackage worksheet from a ScanPipe scan often has duplicate entries for the same purl. I suspect that the number of rows for such a purl reflects the number of resources for that package in the codebaseresources worksheet. In most cases these are exact duplicates, but in some cases one record may have more information (e.g. Description and some additional URLs for bug_tracking, vcs or similar. We should consider deduplicating so that we have one record per purl (preferably the one with the most complete data).
We might want to provide a count of the number of records that have been deduplicated, but I am not sure how important that is.
The discoveredpackage worksheet from a ScanPipe scan often has duplicate entries for the same purl. I suspect that the number of rows for such a purl reflects the number of resources for that package in the codebaseresources worksheet. In most cases these are exact duplicates, but in some cases one record may have more information (e.g. Description and some additional URLs for bug_tracking, vcs or similar. We should consider deduplicating so that we have one record per purl (preferably the one with the most complete data).
We might want to provide a count of the number of records that have been deduplicated, but I am not sure how important that is.