Difference between revisions of "LangID"

From truxwiki.com
Jump to navigation Jump to search
(Created page with "{| style="float:right;border:1px solid black" |+ Details | Executable | <code>LangID.exe</code> |- | Stage | style="text-align:center;" | 18 |- | Percent Complete | style="tex...")
 
Line 38: Line 38:
  
 
Note: The types [[Type_Extracted_ASCII]] and [[Type_Extracted_Unicode]] are typically provided by the [[TextExtract]] [[ETL]].
 
Note: The types [[Type_Extracted_ASCII]] and [[Type_Extracted_Unicode]] are typically provided by the [[TextExtract]] [[ETL]].
 +
 +
=Configuration=
 +
 +
{| class="wikitable"
 +
|-
 +
! Name
 +
! Data Type
 +
! Description
 +
|-
 +
| langid_min_probability
 +
| float
 +
| The minimum threshold for a positive language identification. Between 0.0 and 1.0. Default is 0.333.
 +
|-
 +
| langid_max_bytes_to_interpret
 +
| integer
 +
| The maximum bytes to scan for language identification. Default is 1000.
 +
|}

Revision as of 11:12, 22 June 2021

Details
Executable LangID.exe
Stage 18
Percent Complete 100%
Message Queue langid

The LangID ETL scans ascii and unicode files and identifies non-english language. If it is over a certain confidence threshold it will then tag the document with the language it found.

File Types

LangID processes the following types of files.

File Type Produces
Type_ASCII_Text Tags
Type_UTF8_Encoded_Text Tags
Type_Extracted_ASCII Tags
Type_Extracted_Unicode Tags

Note: The types Type_Extracted_ASCII and Type_Extracted_Unicode are typically provided by the TextExtract ETL.

Configuration

Name Data Type Description
langid_min_probability float The minimum threshold for a positive language identification. Between 0.0 and 1.0. Default is 0.333.
langid_max_bytes_to_interpret integer The maximum bytes to scan for language identification. Default is 1000.