OpenNLP chunker Biomed data

Our chunker data is derived from the GENIA treebank corpus. However, this corpus has complete nested constituencies instead of just chunks. So we use an algorithm to create the chunks out of the treebank. For this there are currently two algorithms in the jcore-base version of the opennlp chunker. I think the newer one works better than the old one but it is still not perfect.
Now I found these data in our internal file system: /archives/alumni_homes/tomanek/coling/corpora/Genia/chunks/genia_new.chunks.gz

This appears to be the GENIA conversion used originally within the JULIE Lab. We should do crossevaluations on both corpora to see if there tagging differences and also just a plain comparison. Perhaps the old data is better.

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

OpenNLP chunker Biomed data #10

Metadata

Assignees

Labels

Type

Projects

Milestone

Relationships

Development

OpenNLP chunker Biomed data #10

Description

Metadata

Metadata

Assignees

Labels

Type

Projects

Milestone

Relationships

Development

Issue actions