Difference between revisions of "LangID"
Jump to navigation
Jump to search
| Line 55: | Line 55: | ||
| The maximum bytes to scan for language identification. Default is 1000. | | The maximum bytes to scan for language identification. Default is 1000. | ||
|} | |} | ||
| + | |||
| + | E.G. for TruxtonSettings.xml | ||
| + | <source lang="xml"> | ||
| + | <truxton_options> | ||
| + | <!-- ... | ||
| + | other configs | ||
| + | ... | ||
| + | --> | ||
| + | <langid_min_probability>0.75</langid_min_probability> | ||
| + | <langid_max_bytes_to_interpret>12000</langid_max_bytes_to_interpret> | ||
| + | </truxton_options> | ||
| + | </source> | ||
Revision as of 11:13, 22 June 2021
| Executable | LangID.exe
|
| Stage | 18 |
| Percent Complete | 100% |
| Message Queue | langid
|
The LangID ETL scans ascii and unicode files and identifies non-english language. If it is over a certain confidence threshold it will then tag the document with the language it found.
File Types
LangID processes the following types of files.
| File Type | Produces |
|---|---|
| Type_ASCII_Text | Tags |
| Type_UTF8_Encoded_Text | Tags |
| Type_Extracted_ASCII | Tags |
| Type_Extracted_Unicode | Tags |
Note: The types Type_Extracted_ASCII and Type_Extracted_Unicode are typically provided by the TextExtract ETL.
Configuration
| Name | Data Type | Description |
|---|---|---|
| langid_min_probability | float | The minimum threshold for a positive language identification. Between 0.0 and 1.0. Default is 0.333. |
| langid_max_bytes_to_interpret | integer | The maximum bytes to scan for language identification. Default is 1000. |
E.G. for TruxtonSettings.xml
<truxton_options>
<!-- ...
other configs
...
-->
<langid_min_probability>0.75</langid_min_probability>
<langid_max_bytes_to_interpret>12000</langid_max_bytes_to_interpret>
</truxton_options>