SOLR

From truxwiki.com
Revision as of 09:45, 2 April 2021 by Pmuessig (talk | contribs)
Jump to navigation Jump to search
Details
Executable SOLR.exe
Stage 192
Percent Complete 90%
Message Queue solrcontentstage

The SOLR ETL (aka SOLR Contents Indexer) indexes media contents metadata items like artifacts, messages and locations into a separately running SOLR Server. It notably does NOT index file metadata, hashes or file contents (those operations are done in SOLRFile).

It occurs near the end of the load once most ETL operations are complete.

Configuration

Configurations for this ETL are nested under the root

Name Data Type Description
solr_url string The url to the solr server and core to use
solr_parallel_extractions number The number of extractions to do in parallel. Optional. Defaults to 1.
solr_parallel_additions string The number of addition batches to do in parallel. Optional. Defaults to 1.

When tweaking the number of parallel extractions and addition is is recommended to not exceed the number of cores available on your exploitation machine. Additional, each parallel operation will consume additional ram in the solr instance running, so make sure you allocate enough for the JVM or you might run into out of memory issues.

E.G. for TruxtonSettings.xml

<truxton_options>
  <!-- ...
       other configs
       ...
    -->
  <solr_url>http://localhost:8983/solr/truxton-core</solr_url>
  <solr_parallel_extractions>4</solr_parallel_extractions>
  <solr_parallel_additions>2</solr_parallel_additions>
</truxton_options>