Skip to content

ScanCode fails to scan large unrar test file #712

Description

@pombredanne

Reported by @msrb in fabric8-analytics/fabric8-analytics-worker#172

The culprit is in the nulls.txt file in https://pypi.python.org/pypi/py-unrar2/0.99.6 : this is a ~100MB test text file with zeroes that is the thing that ScanCode chokes on. Actually this file contains 100M times the '0' character.

$ python -c "print len(set(open('nulls.txt', 'rb').read()))"
1

There are a couple ways to deal with this that come to mind:

  • only scan up to a certain size for large files (configurable with an option) and return a warning
  • compute the number of unique characters in a file (or start of a file) and use this with a threshold to classify the file as data. Do not scan "data" files. This could be refined by computing the entropy.
  • if the file is "text" and the size above a certain threshold, classify the file as data. Do not scan "data" files.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions