Skip to main content

sbx-swe-lemmatization-sparv-schlyter_soderwall_fsv

Analysis citation Information

Språkbanken (2021). sbx-swe-lemmatization-sparv-schlyter_soderwall_fsv (updated: 2021-06-01). [Analysis]. Enriched and distributed by Språkbanken. https://doi.org/10.23695/symc-qb95
BibTeX Additional ways to cite the dataset.
Full-form lookup of citation forms (lemmas) using Schlyter and Soderwall

Lookup of base forms for Old Swedish using Schlyter and Soderwall. Since MSD is not used in this analysis, candidates are ranked with general precision, and results are based on the token form plus optional spelling variants.

Example

This analysis is used with Sparv. Check out Sparv's quick start guide to get started!

To use this analysis, add the following line under export.annotations in the Sparv corpus configuration file:

- <token>:hist.baseform  # Fallback baseforms from Dalin or Swedberg

In order to use this annotation you need to add the following setting to your Sparv corpus configuration file:

metadata:
  language: swe
  variety: fsv

For more info on how to use Sparv, check out the Sparv documentation.

Example output:

<token baseform="|svaka|svarlika|sværia|">Swerikis</token>
<token baseform="|rike|riki|">Rike</token>
<token baseform="|ar|æ|">ær</token>
<token baseform="|af|">af</token>
<token baseform="|hedhne|hesa|hesna|heþna|">hedne</token>
<token baseform="|værld|vald|værald|">værld</token>
<token baseform="|saman|sami|soma|somi|">saman</token>
<token baseform="|ko|komin|min|mit|">komith</token>
<token baseform="|">,</token>
<token baseform="|af|">af</token>
<token baseform="|svear|svet|sveta|sveti|">swea</token>
<token baseform="|akh|ok|uk|">och</token>
<token baseform="|godha|got|">gotha</token>
<token baseform="|ladh|land|lidh|">landh</token>
<token baseform="|">.</token>

Other references

Type

  • Analysis

Task

  • lemmatization

Unit

  • token

Licence

MIT

Keyword

  • schlyter
  • soderwall

Created

2012-10-23

Updated

2021-06-01

Contact

sb-info@svenska.gu.se