Reported by @msrb in fabric8-analytics/fabric8-analytics-worker#172
The culprit is in the nulls.txt file in https://pypi.python.org/pypi/py-unrar2/0.99.6 : this is a ~100MB test text file with zeroes that is the thing that ScanCode chokes on. Actually this file contains 100M times the '0' character.
$ python -c "print len(set(open('nulls.txt', 'rb').read()))"
1
There are a couple ways to deal with this that come to mind:
- only scan up to a certain size for large files (configurable with an option) and return a warning
- compute the number of unique characters in a file (or start of a file) and use this with a threshold to classify the file as data. Do not scan "data" files. This could be refined by computing the entropy.
- if the file is "text" and the size above a certain threshold, classify the file as data. Do not scan "data" files.
Reported by @msrb in fabric8-analytics/fabric8-analytics-worker#172
The culprit is in the
nulls.txtfile in https://pypi.python.org/pypi/py-unrar2/0.99.6 : this is a ~100MB test text file with zeroes that is the thing that ScanCode chokes on. Actually this file contains 100M times the '0' character.There are a couple ways to deal with this that come to mind: