This idea of this plugin is to help remove license boilerplate from source code and replace it with an SPDX-License-Identifer expression tag.
This could be called espeedixfy or some better name.
It would be a postscan plugin requiring license detection. Then if there is a clear, 100% match license all matched to SPDX ids (for now), it would remove the matched text from the file and replaces it smartly with an SPDX-License-Identifier handling (by lexing the code) the proper comment style to inject this.
Anyone could run it on their code, and submit patches to clean up the boilerplate e.g something that makes it easy on maintainers to clean the stuff at their own pace. Linux kernel maintainers have expressed interest for this:
a tool/script like that would be wonderful.
As an option it could also normalize the copyright statements and create a license text file for the doc if it does not exists there.
There is not much left to get there as scancode spits the exact matched text and lines. So from that we can remove the lines from the scanned source code and determine the comment style then inject the license id , possibly using Pygments to determine and preserve the comment style
This idea of this plugin is to help remove license boilerplate from source code and replace it with an SPDX-License-Identifer expression tag.
This could be called espeedixfy or some better name.
It would be a postscan plugin requiring license detection. Then if there is a clear, 100% match license all matched to SPDX ids (for now), it would remove the matched text from the file and replaces it smartly with an SPDX-License-Identifier handling (by lexing the code) the proper comment style to inject this.
Anyone could run it on their code, and submit patches to clean up the boilerplate e.g something that makes it easy on maintainers to clean the stuff at their own pace. Linux kernel maintainers have expressed interest for this:
As an option it could also normalize the copyright statements and create a license text file for the doc if it does not exists there.
There is not much left to get there as scancode spits the exact matched text and lines. So from that we can remove the lines from the scanned source code and determine the comment style then inject the license id , possibly using Pygments to determine and preserve the comment style