GNOME Dashboard Text Indexer
Jim Krehl <jimmuhk@yahoo.com>


This document is intended to explain the architecture of the indexer
and its api.  See also the toplevel README for the Dashboard.

Overview
--------

The indexer is based on a simple vector based document retrieval
system.  There are three main components:

	* Source Database

		This maintains a list of URIs and metadata associated with
		the file such as mtime or web page title.

	* Vector Database

		This maintains the vector representations of all the
		documents in the index.

	* Basis Manager

		This maintains all the words which the indexer searches for
		in documents to represent them in vector form.

Retrieval operations happen only within a specified namespace.  E.g.
an html backend which indexes web pages will only use a segment of the
index store dedicated to web pages and as such will only be able to
retrieve web pages.  However, the basis of the index is maintained
centrally and is shared by all namespaces.



                    +- - - - - - - - - - - - - - - - - - - - - - -+
                    |                                             |
                                       Namespace A
                    |                                             |
                      +------------------+
                    | | Backend/Frontend |                        |
                      +------------------+
                    |         |                                   |
+--------------+      +------------------+   +----------------+
| BasisManager |----|-|   IndexManager   |---| SourcesManager |   |
+--------------+      +------------------+   +----------------+
       |            |         |                                   |
       |              +------------------+
       |            | |   VectorManager  |                        |
       |              +------------------+
       |            |                                             |
       |            +- - - - - - - - - - - - - - - - - - - - - - -+
       |
       |            +- - - - - - - - - - - - - - - - - - - - - - -+
       |            |                                             |
       |                               Namespace B
       |            |                                             |
       |              +------------------+
       |            | | Backend/Frontend |                        |
       |              +------------------+
       |            |         |                                   |
       |              +------------------+   +----------------+
       +------------|-|   IndexManager   |---| SourcesManager |   |
       |              +------------------+   +----------------+
       |            |         |                                   |
       |              +------------------+
       |            | |   VectorManager  |                        |
       |              +------------------+
       |            |                                             |
       |            +- - - - - - - - - - - - - - - - - - - - - - -+
       |
       .                                 .
       .                                 .
       .                                 .


Instantiating an Indexer
------------------------

Indexers are instantiated as an IndexManager object.  Each namespace
should have one and only one IndexManager object.

	public IndexManager (string store_name)

The store_name refers to the namespace.


Using an Indexer
----------------

There are two primary modes of the indexer:  retrieval and indexing.
Retrieval simply queries the index for matching documents.  Indexing
adds a document to the index.  As such, there are 3 ways to call the
indexer:

	public ArrayList Retrieve (TextReader  text,
	                           string      mime_type,
	                       out ArrayList   extracted_clues)

	public void Index (TextReader  text,
	                   string      uri,
	                   string      mime_type,
	               out ArrayList   extracted_clues)

	public ArrayList IndexAndRetrieve (TextReader  text,
	                                   string      uri,
	                                   string      mime_type,
	                               out ArrayList   extracted_clues)

The parameters:

	TextReader text
		this the actual text to indexed/retrieved on.  the invoking
		scope should know what the encoding of the stream is.

	string mime_type
		this is the mime type of the text (currently only
		"text/plain" and "text/html" are supported)

	string uri
		the identifier of the text to be indexed.  should have the
		"file://" prefix for local files.

	out ArrayList extracted_clues
		this parameter will contain any extracted clues from the
		text, e.g. emails, urls.

	return ArrayList matches
		the Retrieve and IndexAndRetrieve methods return the retrieved
		documents


Indexer Architecture
--------------------

Here's a more detailed view of the document retrieval specific
architecture of the indexer.  The indexer is broken up in to various
modules.  They are:

	* Extractors

		Extractors are objects which extend the System.IO.TextReader
		class.  They sit on top of the incoming Document stream and
		snoop every character.  They are specifically designed to
		find particular clues.  E.g. there is currently a
		URLExtractor which grabs url strings from the stream.
		Extractors are passive, in that they do not affect the
		underlying stream at all and as such a stream can be piped
		through any number of extractors without affecting the
		indexing process.  The list of extracted clues can be
		grabbed from the Extractor instance (it's a property) once
		the stream has been processed.

	* TextFilters

		TextFilters also extend System.IO.TextReader.  They work
		like Extractors in that the document stream can be piped
		through them.  However, TextFilters will strip out unwanted
		characters from the stream.  E.g. currently there's an HTML
		filter which strips html tags from the document.

	* Tokenizer

		The tokenizer takes the filtered stream and extracts the
		word tokens from it.

	* Token Filter Bank

		These filter through the token list and filters out common
		words, reduces words to stems, etc.

	* BasisManager

		The BasisManager maintains the list of recognized tokens and
		reduces the incoming documents to a vector based
		representation of those tokens.  It updates the recognized
		token list with new incoming tokens.

	* VectorManager

		The VectorManager maintains the vector representations of
		all the documents contained in the index.

	* SourcesManager

		The SourcesManager maintains the URIs and certain metadata
		associated with a document, e.g. mtimes.  It can also
		potentially store the extracted clues for searching based on
		more deductive strategies.



+----------------------------+
|                            |
| Dashboard Backend/Frontend | <------------------------------------+
|                            |                                      |
+----------------------------+                                      |
              |                                                     |
              |  document                                           |
              V                                                     |
+----------------------------+              +-------------+         |
|                            |    clues     |             |         |
|         Extractors         | -----------> | MetaData DB |         |
|                            |              |             |         |
+----------------------------+              +-------------+         |
              |                                    ^                |
              |  text stream                       |                |
              V                                    |                |
+----------------------------+                     |                |
|                            |                     +---------+      |
|        Text Filters        |                               |      |
|                            |                               |      |
+----------------------------+                               |      |
              |                                              |      |
              |  text stream                                 |      |
              V                                              |      |
+----------------------------+                               |      |
|                            |                               |      |
|         Tokenizer          |                               |      |
|                            |                               |      |
+----------------------------+                               |      |
              |                                              |      |
              |  tokens                                      |      |
              V                                              |      |
+----------------------------+                               |      |
|                            |                               |      |
|       Token Filters        |                               |      |
|                            |                               |      |
+----------------------------+                               |      |
              |                                              |      |
              |  tokens                                      |      |
              V                                              |      |
+----------------------------+              +-------------+  |      |
|                            |    basis     |             |  |      |
|       Basis Manager        | -----------> |  Basis DB   |  |      |
|                            |              |             |  |      |
+----------------------------+              +-------------+  |      |
              |                   uri                        |      |
              |  vector    +---------------------------------+      |
              V            |                                        |
+----------------------------+              +-------------+         |
|                            |   vector     |             |         |
|       Vector Manager       | -----------> |  Index DB   |         |
|                            |              |             |         |
+----------------------------+              +-------------+         |
              |                  matches                            |
              +-----------------------------------------------------+
