

  
  

  
                               
                               
                               
                               
                         DAFS Library
                               
               Programmer's Guide and Reference
                               
                               
                               
                               
                       Release 1.0alpha0
                               
                               
                               
                               
                               
                               
                               
                               
                               
                            11/9/95
                     RAF Technology, Inc.
                16650 NE 79th Street, Suite 200
                       Redmond, WA 98052
                      Tel: (206) 867-0700
                      FAX: (206) 882-7370
                    Email: illcomm@raf.com
                               
                               
                               
                               
                      _Table of Contents
                               

Table of Contents                                          i

Table of Examples                                          v

Table of Figures                                         vii

Overview                                                   1
 Changes from Release 0.9.2                               1
 Applications                                             2
 Typesetting Conventions                                  2

Programmer's Guide                                         4
 Basic DAFS Concepts                                      4
 Data Structures                                          4
 Unicode                                                  5
 Coding Conventions                                       6
 Header Files                                             7
 Error Messages                                           7
 Entities                                                 7
 Creating and Disposing Entities                          9
 Entity Content                                          10
 Comments                                                14
 Properties                                              14
 Entity Traversal                                        20
 Rearranging the Hierarchy                               21
 Or and And Sub-entities                                 21
 File I/O                                                23
 Entity Types                                            24
 How Entities Refer To Images                            24
 Images                                                  28
 Sorting                                                 33
 Call Back Mechanism                                     33
 DAFS Low-level I/O Routines                             34

Programmer's Reference                                    36
 Initialization                                          36
 Entities                                                36
 Entity Type                                             38
 Interpreting Child Entities                             39
 Entity Transversal                                      39
 Iterative Searching                                     40
 Entity Hierarchy Information                            40
 Rearranging the Entity Hierarchy                        41
 Setting Properties                                      42
 Getting Properties                                      43
 Disposing of Properties                                 44
 Traversing and Finding Properties                       45
 Property Attributes                                     45
 Transferring And Copying Properties                     46
 Comments                                                46
 Contents                                                47
 Fixed Point Numbers                                     49
 Images                                                  49
 Using Images With Entities                              52
 Converting and Manipulating Images                      52
 Image Attributes                                        53
 Entity Bounds                                           55
 Orthogons                                               57
 Boxes                                                   59
 Translating Bounds                                      61
 Unicode Characters                                      62
 Unicode Strings                                         63
 DAFS and Text                                           64
 Override FILE I/O                                       66
 Text Matching                                           67
   Background on Regular Expressions                     69
 Call Backs                                              70
 Byte Swapping                                           71
 Error Messages                                          72
 DAFS-BINARY Utility Routines                            72

Index of Functions and Macros                             75

Appendix A: Trademarks, Copyrights and Acknowledgments    78
                      _Table of Examples
                               

Example 1. Creating an Entity Hierarchy                   10

Example 2. Setting a Glyph Value                          11

Example 3. Setting a Text Value                           11

Example 4. Add a Character                                12

Example 5. Delete a Character                             12

Example 6. Setting a Posset                               13

Example 7. Swapping a UserData Property                   16

Example 8. Property Allocation Shortcut                   17

Example 9. Flag Property                                  18

Example 10. Type Properties                               19

Example 11. Traversing Up the Hierarchy                   20

Example 12. Recursive Traversal                           20

Example 13. Iterative Traversal                           21

Example 14. Or Entities                                   22

Example 15. Area Formats                                  27

Example 16. Run-length Image Format                       29

Example 17. Storing Images in a Separate File             31

Example 18. Callbacks                                     34
                       _Table of Figures
                               

Figure 1. Sample Image                                     8

Figure 2. Example Document Structure                       9

Figure 3. And and Or Entities                             23

Figure 4. Different Ways to Represent Area                25
                           _Overview
                               
While many formats exist for composing a document from
electronic storage onto paper, no satisfactory standard exists
for the reverse process. The Document Attribute Format
Specification (DAFS ) has been designed from its inception to
be a standard for document decomposition, with document image
understanding, OCR, and document interchange in mind. The
design of DAFS was a joint effort by RAF Technology and a
working group organized by the Advanced Research Project
Agency (ARPA). The format has been implemented by RAF
Technology. The key features of DAFS are hierarchical
structure, provision for expressing ambiguity and uncertainty,
the ability to encode all languages and scripts, and uniform
treatment and file storage of image and text.  The DAFS
Library, or DAFSLib, is the C-language API (Application
Programmer's Interface) that allows developers to read, write,
construct, and manipulate DAFS files.

This manual is designed for the application programmer who
wishes to program using the DAFS library, or for anyone who
wants a more in depth understanding of the DAFS format. It has
two major sections. The first section is the Programmer's
Guide, and covers the basic concepts of the DAFS file format
and the DAFS Library. The second section is the Programmer's
Reference, and has complete information on all public DAFS
data structures and functions.

DAFSLib is available in a compiled form for Sun SPARC
workstations, SGI Workstations, and PCs running Linux (a free
Unix clone). The library has no user interface, but is usually
shipped with illum (also known as IlluminatorIlluminator
(executable file illum), an X-windows/Motif application for
viewing and editing DAFS files.

Changes from Release 0.9.2

The "borrow entity" concept has been removed from DAFSLib's
repertoire. i_BorrowEntity was intended as a memory saving
device, where the image and text of an Or-type entity could be
borrowed by another Or entity via a pointer to the data. This
allowed possibility sets to be built without duplicating data
in memory. In practice, borrowing was rarely necessary, yet
caused complications in memory apportionment and clean-up. In
most applications, possibility sets are likely only to occur
only with Glyph entities, which under DAFSLib are already a
special case and did not use the borrow feature.

An i_CopyEntity routine has been added in place of
i_BorrowEntity, for cases in which two entities use the same
data.

Applications

Illuminator is one DAFS application mentioned frequently in
this manual, because many DAFSLib users are also Illuminator
users. Illuminator is an editor created for building document
understanding test and training sets, for easy correction of
OCR errors, and for reverse-encoding the text and structure of
a document image. It is configured to display both text and
image, handling most major European languages and Japanese.
Its GUI offers five modes of operation, each of which provides
an essential way to view and edit the data. Please see the
Illuminator User's Manual.

There are many ways that this library could be used DAFS
Library can be used in a wide variety applications, some
examples of which follow.

*  An OCR package could use itDAFSLib to read in a TIFF image
  and write the recognized data in DAFS format.
  
*  A database of accurately labelled images and text (ground
  truth"ground truth") could be distributed in the DAFS
  format. The University of Washington has producedused
  DAFSLib and the Illuminator program to produce such a
  database of English and Japanese text documents.
  
*  An OCR developer might write a character segmenter that
  wrote the image of each character into a DAFS file
  consisting of alltaining only other images of the same
  glyph. The files of sorted images could then be used to
  train an OCR engine. The illum program that comes with
  DAFSIlluminator is designed to aid in construction of such
  databases.
  
*  Document structure recognizers mightcould output their
  information in a DAFS file using the DAFSLib. Many
  commercial OCR formats do not have the flexibility of DAFS
  in representing complicated document structure.
  
Typesetting Conventions

The stylistic conventions used in this manual are summarized
below.

In this manual, there are some stylistic conventions.Quoted C
code is indented and displayed in a typewriter font, thus as
follows:

  int main(int argc, char **argv)
  {
     printf("hello world.\n");
  }
     
A typewriter font will also be used for references to C code
in the narrative, for example text, e. g. "C programs start in
the main function".

Unix programs, pathnames, arguments, etc., will be referred to
in italics will be italicized, e. g. illum foo.dr. Italics
will also be used to introduce new terms.

Type names are referred to withshown in bold letters, e.g.
Word, Zone. Bold is also used in section titles, table column
names, and for emphasis.

Property names are shown in quotes in a typewriter font, just
as they would be supplied as arguments to a C function. For
example, "flag", "bold".

                               
                               
                               
                               
                      Programmer's Guide
                               
Basic DAFS Concepts

The standard unit of DAFS is the entity. Entities are arranged
in a hierarchical structure. Each entity has a number of
standard attributes: a pointer to image, a structure
describing its area on the image, text content, and a type.
Entities can also can have arbitrary properties that the
developer can specify. A property is a name, value pair.
Another important data structure is the DAFS representation of
image, which uses an efficient run-length encoding for image
data.

The entity type of an entity is simply a string associated
with thata particular entity. (Similarly, the property type is
a string associated with a property.) In some other
programming systems, the word "type" implies a regularity and
definite nature that is not present in DAFS. All DAFS entities
are alike; the meaning associated with the entity type is
merely conventional. There are two exceptionsis an exception
to this principle: first, properties may be associated with
types as well as individual entities, second, the illum
program haIlluminator's Text Mode uses some hardwired
assumptions about type names in its Text Mode.

A typical convention might be to give the entities of a
document the following types: Glyph, Space, Word, Line,
Newline, Paragraph, and the Document as a whole. Each entity
is a region of image delimited by an image bounds. For
example, a Word entity is represented by a region bounding the
image of a word on screen. Typically, the image of a line of
text will be associated with a Line entity which may have Word
child entities; each Word may have Glyph children; each Glyph
may be labeled with its textual content (a character). The
choice of names and the structure of the hierarchy in the
above example is purely conventional, and the applications
programmer is free to invent his own conventions.

Unicode, the emerging universal character encoding standard,
allows DAFS to handle files written in almost any foreign
script. Illuminator can display and edit documents written in
many different languages, including several Japanese scripts.

Currently DAFS files may be stored in DAFS-BINARY format. The
ASCII text portion of a DAFS file (the combined entity textual
content) may also be saved separately as ASCII text with
Unicode characters encoded using the UTF/FSS multi-byte
encoding scheme, or as Unicode text encoded according to the
Unicode Standard. Of course, all the information about the
image or the bounding boxes is lost in such a transformation.
SGML output is planned for the future.

Unicode, the emerging universal character encoding standard,
allows DAFS to handle files written in almost any foreign
script.

Data Structures

The following are opaque definitions of the three main data
types in DAFS: the entity, the image and the property.

  typedef struct iEntity *iEntityPtr;
  typedef struct iImage *iImagePtr;
  typedef struct iProp *iPropPtr;
     
The most common form for describing areas in DAFS is the box:

  typedef struct iBox {
   short x, y, w, h;
  } iBox;
     
An iBox describes a rectangular box, where x is zero on the
left of the image and increases to the right, y is zero at the
top and increases down the page, w is the width, and h is the
height. Other data structures for describing areas are
described lateron page 24.

Another important basic data type in DAFS is the fixed point
number:

  typedef long FixedPoint;
     
By convention, the high two bytes of a Fixed Point number is
an integer with a value naturally ranging between [0  216),
and the low 2 bytes represents a fraction part numerated in
parts of 2-16. Such a format has a number of useful
properties. Although it cannot represent the range of values
that a floating point number can, it can be converted to and
from an integer with a shift, and it can dundergo shift,
addition and subtraction operatorions as easily and as
fastquickly as a normal integer. Multiplication or division
require an extra shift. The library supplies convenient macros
for manipulating fixed point numbers. This data type is used
to specify magnification levels for images.

Unicode

DAFSLib uses two character encoding standards, ASCII and
Unicode. Unicode covers almost all the world's scripts. (The
few remaining are either ancient or obscure, and will likely
be included in future revisions). Most text processing
functions come in both flavorsASCII can encode primarily the
European language scripts, using one byte per character.
Unicode uses two bytes per character, whereas ASCII is one-
byteand covers almost all the world's language scripts. Most
text processing functions are capable of using both. DAFSLib
supplies a library that imitatesof functions which imitate the
standard string and character functions that operate on ASCII
strings and characters, except they operate on Unicode strings
and characters.

Unicode is a text encoding, and not a text rendering, scheme.
Unlike most extended ASCII's (versions of ASCII that add
national characters to the upper 128 characters in a byte),
Unicode discourages  diacritical mark characters (such as )
and combination characters (those which could just as well be
represented by combining two also discourages combination
character and diacritical mark characters, such as , whichor
more existing Unicode characters). Unicode canonically
represents by ansuch characters by a character followed by one
or  more nonspacing characters, for example, `a' followed by a
grave. In this case, Unicode does have a combination a-with-
grave character for the sake of backwards compatibility, but
this is one of the exceptions.

Unicode for the most part liftborrows entire national
standards, by simply tacking on an extra byte. For example,
the Russian national standard is a single-byte extended ASCII
where the lower 128 characters are ASCII as usual, and the
upper 128 characters are cyrillic. Unicode already has ASCII
at high-byte index "0x00" and places the Cyrillic characters
at high-byte index "0x04". But Unicode also tries to unify
characters that appear in different national character sets.
The most ambitious example of this is Han characters, which
are used in China, Japan and Korea. Each of those countries
has its own standards for the ordering of the characters, but
the actual meaning of the characters largely overlaps. Unicode
has unified those characters into one set. This makes
converting characters from national Han standards a lookup
table affair instead of a simple offset.

ASCII is a subset of Unicode, and by design, all characters in
ASCII are the same in Unicode, but with an added index of
0x00. Most operating systems are at a loss when processing
Unicode characters, and so Ken Thompson of AT&T came up with a
scheme to encode Unicode characters in a multibyte form called
UTF-FSS. These conversion routines are supplied with DAFS. The
nice properties of UTF-FSS are that Unicode characters that
can bare mapped directly to ASCII characters, ar if possible,
and other characters map to 2-4 byte sequences that are
guaranteed not to be mistaken for ASCII characters or special
characters like NUL. Some Unicode programs actually use this
multi-byte ASCII format for external files rather than Unicode
itself.

Coding Conventions

The programmers of DAFS observed the following conventions,
which we hope will aid the developer in understanding the
code.

  iTypeNameOrEnumValue
  i_FunctionName()
  MACRO
  local_variable
     
Most functions refer to the type that they primarily operate
on in their name. For example, i_UncompressImage operates on
an iImagePtr. Whenever a function does not refer to a type in
its name, it usually operates on an entity. For example,
i_GetText retrieves the text value of a given iEntityPtr.
"Prop" instead of "Property" appears in functions operating on
iPropPtr's or properties.

DAFSLib has get and set functions on pointer data. Usually,
the set functions make a copy of the data and stores it in the
DAFSLib structure, whereas the get function returns a pointer
to the stored data that "belongs" to DAFS. To put it another
way, the set function copies its input and so the input
pointer's management is the responsibility of the application
programmer, and the get function returns a pointer whose
management is the responsibility of DAFS. Modifying the data
pointed to by a pointer returned from a get function is
marginally OK, although it short-circuits the callback
mechanism, which we explain later. Freeing, or reallocating
the pointer itself is forbidden.

The examples are somewhat stylized for internal quality
control reasons.

Header Files

In order to use all DAFS functions, it is only necessary only
to include dafs.h. It in turn includes dafstype.h for data
type definitions, ucs.h for Unicode utility functions, and
utf.h for Unicode to ASCII translation. For access to the
internals of the run-length image implementation, one must
include rlimage.h.

Error Messages

Only the simplest DAFS functions return a value. Most return
an error code of type iDAFSError. The definition of DAFS error
codes are subject to revision, but the condition of no error
will always be represented by DAFSOK, which has a value of 0.

The DAFS programming philosophy says that error codes are to
be returned whenever the error is due to the user or to
transient operating system conditions, for example, reading a
file that isn't there, or running out of memory. Basically,
anything that the user can do something about should bubble up
an error. A different class of error is programmer error or an
inconsistent state inside of DAFSLib. DAFSLib really has no
idea what is going on or whohow to proceed in such a case; if
DAFSLib detects such an error, it halts immediately.

Entities

Any object of interest in a document can become an entity. An
entity can have a pointer to an associated image, a bounding
box on that image, text content, and properties, but it need
not have any of these. It also hapossesses an entity type and
has a position in the hierarchy.

For didactic purposesTo illustrate the concepts of entity and
entity hierarchy, we introduce the following image, viewed in
the image mode of the illum programImage Mode of Illuminator:







                               
                               
                    Figure 1. Sample Image
                               


whichThe preceding image has been imperfectly recognized by an
OCR program, resulting in the following entity hierarchy:



                               
                               
                               
             Figure 2. Example Document Structure
                               

Creating and Disposing Entities

i_NewEntity creates a new entity of the type specified and
returns a pointer to it with the default bounds being set to
zero, while i_DisposeEntity frees all memory associated with
an entity and also disposes of all its descendants. There are
several other functions for creating entities. i_CopyEntity
copies the contents of one entity into another.

In order to create the entity hierarchy in our running
example, we could write this fragment of code.

            Example 1. Creating an Entity Hierarchy
  /* begin example: Hierarchy */
  /* we will make these global variables so we can use them in
     later examples */
  iEntityPtr doc;
  iEntityPtr para1;
  iEntityPtr para2;
  iEntityPtr word1, word2, word3;
  iEntityPtr glyph1_1, glyph1_2, glyph2_1, glyph2_2, glyph2_3;
  iEntityPtr glyph3_1, glyph3_2, glyph3_3, glyph3_4;
  
  
  void
  makeHierarchy()
  {
     /* later on we will discuss the manner in which entity
     types are defined */
     doc = i_NewEntity(iDocumentT, 0);    /* supply a null
     parent */
     para1 = i_NewEntity(iParagraphT, doc);
     para2 = i_NewEntity(iParagraphT, doc); /* para1 and para2
     will stay in the order they were created unless the
     hierarchy is rearranged */
     word1 = i_NewEntity(iWordT, para1);
       /* indentation for illustrative purposes only */
       glyph1_1 = i_NewEntity(iGlyphT, word1);
       glyph1_2 = i_NewEntity(iGlyphT, word1);
     word2 = i_NewEntity(iWordT, para1);
       glyph2_1 = i_NewEntity(iGlyphT, word2);
       glyph2_2 = i_NewEntity(iGlyphT, word2);
       glyph2_3 = i_NewEntity(iGlyphT, word2);
     word3 = i_NewEntity(iWordT, para1);
       glyph3_1 = i_NewEntity(iGlyphT, word3);
       glyph3_2 = i_NewEntity(iGlyphT, word3);
       glyph3_3 = i_NewEntity(iGlyphT, word3);
       glyph3_4 = i_NewEntity(iGlyphT, word3);
     /* this example did not show associating text or image
     with an entity */
  }
  /* end example: Hierarchy */
     
Entity Content

An entity can have different types of content. Each entity may
possess an image content, the area of image enclosed by its
bounding box. Each may also possess a textual content. A
character, a string, a set of possible characters or user-
defined data can all be stored as the textual content of an
entity. When used by itself in this manual , content refers to
textual content.

i_GetGlyph and i_SetGlyph are used to save and retrieve a
single character value. To get and put an ASCII string, use
i_GetText and i_SetText. There are also i_GetUText and
i_SetUText for Unicode strings. For both types of string, the
set function allocates space and makes a copy, and the get
function returns a pointer to the copy made at the time of the
set. This data belongs to the entity and the pointer ought not
to be destroyed or modified by the user. In order to clear all
data from an entity, use i_SetNoData.

To give an example, the first character in the our running
example could be set thus:

               Example 2. Setting a Glyph Value
  /* begin example: glyph */
  void setOneGlyph() {
     i_SetGlyph(glyph1_1, 'M', 100);
  }
  /* end example: glyph */
     
To set the text of the second paragraph, we could use

                Example 3. Setting a Text Value
  /* begin example: text */
  void setOneString() {
     i_SetText(para2, "The dog chased the cat.");
  }
  /* end example: text */
     
To edit the text to add a character '.', we could do this

                  Example 4. Add a Character
  /* begin example: addchar */
  void addPeriod() {
     const char *s;
     char *new_s;
     int len;
  
     s = i_GetText(para2);
     len = strlen(s);
  
     /* we must make a copy to modify the value because we
      don't know whether there is enough space
      in the existing pointer */
     new_s = (char *) malloc(len + 2); /* 1 for NUL, 1 for
     character */
     strcpy(new_s, s);
     new_s[len] = '.';
     i_SetText(para2, new_s);
  
     /* we are responsible for freeing this
      since DAFS made a copy */
     free(new_s);
  }
  /* end example: addchar */
     
In order to delete one character from the end of the entity's
text, we could get by with:

                 Example 5. Delete a Character
  /* begin example: delchar */
  void deleteAChar(){
     char *s;
     int len;
  
     s = (char *) i_GetText(para2); /* defeat const */
     len = strlen(s);
  
     s[len] = '\0'; /* we know that we will not get a memory
     error */
     /* leave out the next statement and the entity's text is
     still changed, but callback isn't called */
     i_SetText(para2, s);
  }
  /* end example: delchar */
     
i_SetText makes a copy of its new data before freeing the old,
otherwise, we would free the modified pointer before we used
it.

Later on we will see how to express ambiguity or uncertainty
in any type of entity using the iOr entity concept. Here we
will discuss ambiguous Glyph entities, which are treated as a
special case. When a Glyph's text content might be one of
several character values, the content is represented by a
possibility set. A possibility set is a collection of
characters paired with confidences. In our running example,
the OCR was uncertain of the value of the first character of
the third word (glyph3_1), so it gave it two possiblities,
'a' with a 50 confidence, and 'r' with a 34 confidence).
We could assign this value with the following code.

                  Example 6. Setting a Posset
  /* begin example: posset */
  void setPosset(){
     iPosSet ps[2];
     ps[0].chr = 'a';
     ps[0].conf = 50;
     ps[1].chr = 'r';
     ps[1].conf = 34;
     i_SetPosSet(glyph3_1, 2, ps);
  }
  /* end example: posset */
     
In a manner similar to i_SetText, i_SetPosSet makes a copy of
its input data, and so it is ok that ps is an automatic
variable. The meaning of the attached confidence is set by
convention and is not necessarily a probabilty. i_SetPosSet
and i_GetPosSet can be used to set an entity's content to zero
or to many different possible character values each with a
different confidence. An entity can only have one type of
content data (character or string) at a time, however. If a
program calls i_SetGlyph, setting the entity contents to 'b',
then i_SetText to make the contents "dog" and finally
i_GetGlyph, the value of 'b' will not be returned because the
content of the entity is now a text string. To find out what
type of data is stored with an entity, use i_GetEDataType.

                     Table 1. Text Content
    Textual Content         Information Stored
      Enumeration
          iDNo              There is no textual
                                  content
        iDPosSet                 A set of
                           character/confidence
                                   pairs
        iDString              An ASCII string
       iDUString             A Unicode string
       iD1PosSet                 A single
                           character/confidence
                                   pair
       iDUserData           Arbitrary user data
         iDConf                A confidence
        iDShort             A two byte integer


Comments

Comments can be stored with each entity using i_SetComment,
and i_GetComment will return the entity's comment. To remove a
comment, use i_SetComment with a NULL string.

Properties

Properties are attributes that are associated with an entity
or entity type. Examples of properties are "bold", "font size"
or "italic". Properties can represent a wide variety of
attributes. These attributes can have several different kinds
of values. For instance italic would have a boolean value; a
word can only be either italicized or not. However, an
attribute such as "font size" requires quantification (eg.
14pt); it would be represented by an integer property. Other
property types exist, as the following table shows:

                 Table 2. Property Value Types
   Property type (by         Attribute represented
      enumerator)
        iPBool            a boolean property (ie. true
                                   or false)
        iPLong            a number representable by a
                               four byte integer
       iPString                 an ASCII string
       iPUString                a Unicode string
      iPUserData         pointer to contiguous section
                                   of memory
       iPUserPtr                   a pointer


The last two properties, iPUserData and iPUserPtr, allow for
data types and structures that DAFS library doesn't know
about.

iPUserData is for data that is to be saved with DAFS file. It
can be an array, or a structure, but the structure cannot
itself contain pointers, because a copy is made of the value
and it is written to and read from disk in a naive fashion.
iPUserData also has a problem with byte-swapping. Byte-
swapping is something is necessary whenever data is copied
between different computer architectures that order bytes
differently. For example if a short is located at address
0x1234 on a Motorola 68000 processor, then the more
significant byte of the short will be located at 0x1234 and
the less significant byte will be located at 0x1235, in
contrast, on an Intel 8086 processor, the more significant
byte would be located at 0x1235 and less significant byte at
0x1234. Although DAFSLib automatically takes into account byte-
swapping for all other DAFS data, it does not know how to swap
an iPUserData property, because it does not know what it
really is. Because of this, the application programmer must
supply a swapping function when retrieving such data. DAFSLib
supplies some utility functions for swapping the components of
a structure.

Here is an example of a property and its swapping function

            Example 7. Swapping a UserData Property
  /* begin example: swap */
  struct MyData {
     long a_long;
     short a_short;
  };
  
  void swapMyData(void *p)
  {
     struct MyData *my_data = (struct MyData *) p;
  
     my_data->a_long = i_SwapL(my_data->a_long);
     my_data->a_short = i_SwapS(my_data->a_short);
  }
  
  void testSwapping(){
     const void *my_data;
     iEntityPtr entity;
     long size;
     iDAFSError e;
  
     e = i_GetUserDataProp(entity, "my data", (void *)
     &my_data, &size,    swapMyData);
  }
  /* end example: swap */
     
DAFS knows whether this property was read in from a file, and
the byte order of that file was, or whether it was created
during this session; and it knows what the byte order of this
computer is. It can therefore decide whether to call the
swapping function or not.

One useful trick is to send in a NULL pointer for the data,
and supply a size. Under these conditions, DAFS allocates the
suggested amount of space, and knows not to try and copy the
NULL pointer, so you can then get the allocated space and set
it afterwards. This is a way to avoid doing two allocations,
one on the user side and one on the DAFS side. Here is an
example:

            Example 8. Property Allocation Shortcut
  /* begin example: alloc */
  /* return an array that contains the number of black pixels
     in each horizontal line -- the horizontal sum of pixels
     */
  void calcHsmp(iImagePtr image, iBox b, short *hsmp)
  {
     char *bits;
     int2 data_width;
     iFixedPoint mag = 1L<<16;
     int i, j, k;
     int sum;
  
     i_Image2BitMap(image, b, mag, &data_width, &bits);
     for(i = 0; i < b.h; i++) { /* loop over lines
       sum = 0;
       for(j = 0; j < data_width; j++) /* loop over bytes */
          for(k = 0; k < 8; k++) /* loop over bits */
               sum += ((1 << k) & bits[i*data_width + j]) !=
     0;
       hsmp[i] = sum;
     }
  }
  
  void allocShortCut(){
     short *hsmp;
     iEntityPtr entity;
     iDAFSError e;
     iBox b;
     #define HSMP "HSMP"
  
     b = i_GetEnclosingBox(entity);
     e = i_SetUserDataProp(entity, HSMP,0, sizeof(short)*
     b.h);
     e = (e != DAFSOK) ?
      e :
      i_GetUserDataProp(entity,HSMP,(void**)&hsmp, 0, 0);
     if(e == DAFSOK)
       calcHsmp(i_GetImage(entity), b, hsmp);
  }
  /* end example: alloc */
     
iPUserPtr is for transient data that only exists for the life
of the program. Instead of attempting to copy the memory
pointed to, which could be an arbitrarily complicated C
structure with pointers inside of it, DAFSLib merely remembers
the actual pointer address. If the entity owning the property
is written to disk, iPUserPtr properties are ignored, because
the pointer values are meaningless outside of the context of
the core image of the running program. The application
programmer can either supply a dispose callback to go with the
iPUserPtr property, to be called when the property is freed,
or keep the responsibility for managing the memory for
himself.

Properties can be created and associated with an entity by the
i_SetXXXProp functions, where XXX is the type of data, such as
String, UserData, etc. The i_GetXXXProp functions are
typically used to retrieve properties from an entity. If an
entity doesn't have the specified property, DAFSLib will then
determine if the entity's type has the property and return
that property if it exists. Session properties are discussed
in more detail below.

In our running example, the OCR program came up with two
possible values for one of the characters. It really ought to
flag that entity as something the user should double-check.
Here is some example code:

                   Example 9. Flag Property
  /* begin example: flag */
  void flagIt() {
     i_SetLongProp(glyph3_1, "flag", 1);
  }
  /* end example: flag */
     
It happens that the illum programIlluminator takes note of the
"flag" property and has methods for speedily correcting them.
It is a long property instead of a boolean property so that
the number could perhaps represent the reason for flagging the
entity.

By convention, if a property is not listed for an entity, the
entity was not categorized with respect to that property. For
example, if an entity is not bold it must have the "bold"
property and the "bold" property must be set to False. If the
entity was not categorized with respect to bold then there
should be no "bold" property attached to it. The exception is
the "flag" property we mentioned above.

Properties are given a name when they are created. This name
is an ASCII string that will be used to reference the
property. There are no predefined properties; an application
is free to establish any properties it wishes. There are a few
widely-used conventions: the name should be in lower case, and
one word if possible. The "comment" property is used by
DAFSLib to implement the i_SetComment and i_GetComment
routines. The "illum" and "editing" properties are reserved
for the use of the illum programIlluminator.

DAFS also allows you to associate a property with entities of
a given type, using the i_SetSessionXXXProp functions. Then
when you get a property on an entity of that type it first
looks to see whether the entity has that property and then if
it does not, it checks to see if that property is associated
with the type. Here are some examples:

                  Example 10. Type Properties
  /* begin example: session */
  #define TEXT_DIRECTION "Text Direction"
  #define L_TO_R "left to right"
  #define R_TO_L "right to left"
  
  void illustrateSessionProps(iEntityPtr english_line,
                              iEntityPtr arabic_line)
  {
     char *value;
  
     i_SetSessionStringProp(iLineT, TEXT_DIRECTION, L_TO_R);
     /* in case you had any doubts about my example */
     i_SetEntityType(arabic_line, iLineT);
     i_SetEntityType(english_line, iLineT);
  
     i_SetStringProp(arabic_line,TEXT_DIRECTION, R_TO_L);
  
     i_GetStringProp(english_line, TEXT_DIRECTION, &value);
       /* will return "left to     right" because there is no
     TEXT_DIRECTION property  associated with english_line,
     and so it defaults to the     value associated with the
     line type */
  
     i_GetStringProp(arabic_line, TEXT_DIRECTION, &value);
       /* will return "right to    left" because we have
     overridden the type property with  one local to the
     entity */
  }
  /* end example: session */
     
There are no i_GetSessionXXXProp functions. The properties
that are set with the i_SetSessionXXXProp functions are not
read in or written out to disk, so, at most, they will only
last as long as the application is run.

Note: In previous versions of DAFSLib the properties that were
associated with entity types were written out to and read in
from disk. In this version they are not.

i_MoveAllProps can be used to remove the properties attached
to one entity and reattach them to another. i_CopyAllProps
will add the properties to a different entity, but leave the
original entity unchanged.

Entity Traversal

To move around the entity hierarchy, DAFSLib provides
i_GetNextEntity, i_GetPrevEntity, i_GetChildEntity,
i_GetLastEntity, and i_GetParentEntity.

Here is some example code for finding the root of a hierarchy
given a single entity.

            Example 11. Traversing Up the Hierarchy
  /* begin example: find parent */
  iEntityPtr findRoot(iEntityPtr entity)
  {
     iEntityPtr parent;
     while(parent = i_GetParentEntity(entity))
       entity = parent;
     return entity;
  }
  /* end example find parent */
  
Here is an example of a recursive routine to print out the
values of a portion of the hierarchy.

                Example 12. Recursive Traversal
  /*begin example: recurse */
  void printRecurse(iEntityPtr root)
  {
     iEntityPtr child;
     void printOne(iEntityPtr entity);
  
     printOne(root); /* I will describe in the next example */
  
     for(child = i_GetChildEntity(root);
       child;
       child = i_GetNextEntity(child)){
          printRecurse(child);
     }
  }
  /* end example: recurse */
     
i_FindEntityForward and i_FindEntityBack are for searching in
read order for entities that fit a particular criterion.
Although it is tempting to use them instead of recursive code,
they do not limit themselves to the subtree rooted at the
initial argument. Let us define printOne and printIterative to
demonstrate the difference.

                Example 13. Iterative Traversal
  /* begin example: iterate */
  void printOne(iEntityPtr en)
  {
     iDTag tag;
     short c;
     char *s;
  
     tag = i_GetEDataType(en);
     /* beware, we aren't handling all types of data */
     switch(tag){
       case iD1PosSet:
          c = i_GetGlyph(en, 0);
          printf("%c ", c);
          break;
       case iDString:
          s = i_GetText(en);
          printf("%s ", s);
     }
  }
  
  void printIterative(iEntityPtr en)
  {
     for(; en; en = i_FindEntityForward(en, 0, ALL, 0)){
       printOne(en);
     }
  }
  /* end example: iterate */
     
The output of printRecurse, given the entity word2, would be
"c a t ". The output of printIterative would be begin the same
way, but continue on until the end of the document to produce
"c a t a a n . The dog chased the cat. ". On the other hand,
another advantage of the find functions is that having the
current entity is the same as having the entire state
information to continue the search. One use of this would be
to write a function in a GUI that searched for flagged
entities. The user could search, edit the flagged entity, and
then search again.

Rearranging the Hierarchy

i_MoveToFront, i_MoveToBack, i_MoveBefore and i_MoveAfter will
rearrange the hierarchy of entities.

Or and And Sub-entities

The children on an entity can be interpreted in two different
ways, as or children or as and children, according to the
LogicalType parameter. iOr is used to indicate that all the
children of this entity are different possible alternatives of
the parent entity, implying that just one of these
alternatives must be chosen to reconstruct the entity. By
convention, the types of the iOr children are the same as that
of the parent, however, nothing enforces this convention.
iAnd, in contrast, indicates that the parent is composed of
all the children concatenated. Continuing the read order
example of the previous paragraph, a document recognition
package could create one iOr ReadOrder entity that has three
ReadOrder children, each describing a different possible read
order. The examples of hierarchy we have seen so far are those
of iAnd children, for example an iAnd Word entity that has
several Glyph entity children which together spell the word.
i_GetLogicalType and i_SetLogicalType are used to get and set
the value of the LogicalType parameter.

We give an example based on our running example. We make a
structure that OR and AND entities. The first word in our
example is changed to be this:

                    Example 14. Or Entities
  /* begin example: or */
  void illustrateOr(){
     iEntityPtr word1a, word1b;
     iEntityPtr glyph1_1a, glyph1_1b;
     iEntityPtr glyph1_2b;
     iDAFSError e;
  
     word1a = i_NewEntity(iWordT, word1);
     word1b = i_NewEntity(iWordT, word1);
     if(word1a == 0 || word1b == 0)
       e = MallocFailError;
     /* !e is short for e == DAFSOK */
     if(!e) {
       i_SetLogicalType(word1, iOr);
       i_MoveToBack(glyph1_1, word1a);
       i_MoveToBack(glyph1_2, word1a);
       glyph1_1a = i_NewEntity(iGlyphT, word1b);
       glyph1_1b = i_NewEntity(iGlyphT, word1b);
       if((glyph1_1a == 0 || glyph1_1b == 0))
          e = MallocFailError;
     }
     if(!e) {
       i_SetGlyph(glyph1_1a, 'T', 0);
       i_SetGlyph(glyph1_1b, 'h', 0);
     }
     if(!e) e = i_CopyEntity(glyph1_2, word1b, &glyph1_2b);
  }
  /* end example: or */
     
The result of this is the following structure:

                               
                               
                 Figure 3. And and Or Entities
File I/O

i_ReadEntity is used to read the contents of a file into a new
entity, returning a pointer to it. i_WriteEntity writes an
entity and the entire hierarchy under it to a file.

Some sample code for reading:

  /* begin example: read */
  void readIt() {
     iEntityPtr root;
     iDAFSError e;
     e = i_ReadEntity("foo.dr", &root);
  }
  /* end example: read */
     
Some sample code for writing:

  /* begin example: write */
  void writeIt() {
     iDAFSError e;
     e = i_WriteEntity("dogs_and_cats.dr", doc, DAFSB);
  }
  /* end example: write */
     
It is possible to write the entity structure to one file and
the images to other files. i_SetImageCallBack is used to pass
a callback routine to DAFSLib that will be called whenever
DAFSLib wants to write an image or to read an image that was
written with the write callback. To identify the image to the
callback routines i_SetImageTag and i_GetImageTag can be used.
We have an example of this later on this section.

Entity Types

The type of an entity is set when the entity is created. It
can be retrieved and changed with i_GetEntityType and
i_SetEntityType. Types are implemented as ordinal numbers, but
also have associated strings, and the conversion is
accomplished via i_NameToType and i_TypeToName. There are some
pre-defined types that are represented by convenient enum
values. In the first example, we used some of these:
iDocumentT, iParagraphT, etc. The index associated with a type
can vary from file to file, but it is constant during an
instance of running a program, and the index in a disk file is
rationalized with the index in the core image of the library
when the file is read in. Properties can also be attached to
each entity type, as discussed in the section on properties.

How Entities Refer To Images

It is usual for all the images in a hierarchy to refer to the
same image. In fact, whenever an entity is created, it is
automatically set to point to the same image as its parent
does. (An entity hierarchy can still refer to multiple images
or none at all). The library keeps track of how many entities
are using each image, so that it can do automatic garbage
collection on images that are no longer referred to. The
entity's image is set with i_SetImage and retrieved with
i_GetImage.

Each entity can refer to a portion of its image using one of
three formats. The bounds for the portion's boundary is
obtained from i_GetArea. The routine will return the bounds
for the image portion, along with the format the bounds are
in. The three formats are:

 iABox - a simple rectangular shape
 iABoxList - a collection of rectangles
 iAOrthogonList - a collection of orthogons

An orthogon is a polygon where the angles between successive
points are right angles. The i_SetBox, i_SetBoxList, and
i_SetOrthogonList are used to assign areas of the appropriate
type. i_GetEnclosingBox returns the box that encloses the
area, no matter what its internal representation, without
changing that representation.

                               
                               
          Figure 4. Different Ways to Represent Area
Because the orthogon and box list formats use pointers, the
structures need to be initialized and disposed of. Functions
for the box list format include i_NewBoxList,
i_DisposeBoxList, and i_CopyBoxList. Similar functions for the
orthogon list format are i_NewOrthogonList,
i_DisposeOrthogonList, and i_CopyOrthogonList. There are also
functions for individual orthogons: i_NewOrthogon,
i_CopyOrthogon, and i_DisposeOrthogon. i_DisposeArea will
remove the image bounds from an entity regardless of its
format.

Each image bounds format has some functions to determine how
and if two of them overlap: i_IntersectBox and i_UnionBox;
i_IntersectOrthogon and i_UnionOrthogon; i
_IntersectOrthogonList, i_UnionOrthogonList, and
i_SubtractOrthogonList; and i_UnionBoxList,
i_IntersectBoxList, and i_SubtractBoxList.

i_ForceAreaType will convert an entity's image bounds format
to the specified type. Other conversion functions are
i_BoxList2OrthogonList, i_OrthogonList2Box, i_Box2BoxList,
i_BoxList2Box, i_Box2OrthogonList and i_OrthogonList2BoxList.
Conversions to iBox are possibly lossy, by which we mean that
it can only represent the enclosing box of the more
complicated areas described in the other formats. The current
implementation of iOrthogonList is not capable of representing
regions with holes in them.

The box list format does the intersection and union operations
efficiently. The orthogon format is convenient for rendering
the regions in a windowing system.

In the following example, we show how to interpret the various
formats.

                   Example 15. Area Formats
  /* begin example: area */
  void printArea(iEntityPtr entity)
  {
     void *data;
     iATag type;
     iBox box;
     iBoxListPtr box_list;
     iOrthogonListPtr ortho_list;
     iOrthogonPtr ortho;
     int i;
     int j;
  
     i_GetArea(entity, &data, &type);
     switch(type){
     case iABox:
       box = *(iBox *)data;
       printf("box: (%d, %d) %dx%d\n",
         box.x, box.y,
         box.w, box.h);
       break;
     case iABoxList:
       box_list = (iBoxListPtr) data;
       printf("box_list: enclosing box (%d, %d) to (%d,
     %d)\n",
          box_list->extents.x1, box_list->extents.y1,
          box_list->extents.x2, box_list->extents.y2);
       for(i = 0; i < box_list->numRects; i++){
          printf("\tbox %d: (%d,%d) to (%d, %d)\n",
          i,
                box_list->rects[i].x1,
                box_list->rects[i].y1,
                box_list->rects[i].x2,
          box_list->rects[i].y2);
       }
       break;
     case iAOrthogonList:
       ortho_list = (iOrthogonListPtr) data;
       printf("orthogon list:\n");
       for(i = 0; i < ortho_list->num_poly; i++){
          ortho = ortho_list->poly[i];
          printf("\torthogon %d:\n", i);
          for(j = 0; j < ortho->n; j++){
               printf("\t\tpoint %d: (%x, %y)\n",
                    j,
                    ortho->point[j].x,
                    ortho->point[j].y);
          }
       }
     break;
    }
  }
  /* end example: area */
     
Images

i_ReadImage will read either a TIFF or a PDA image and return
an iImagePtr that points to it in memory. A file name's
extension, ".pda" for example, determines how the library will
read the image. To check to see if the extension is one that
the library supports, use i_ImageExt. i_WriteImage will write
images to a file in either TIFF or PDA format, and may be
uncompressed, group III or group IV compressed.

An iImagePtr points to a binary image that is stored in memory
in a run-length-compressed format. This format takes about one
third the memory that a bitmap would take on a typical page.
It also provides faster magnification, reading and writing to
the disk. The i_GetImageWidth and i_GetImageHeight will return
the image's width and height. The i_GetImageBounds routine
will return height and width of the image in the form of an
iBox structure (that is, width, height and upper left corner
coordinates). To get bitmaps into and out of an iImagePtr, the
routines i_BitMap2Image and i_Image2BitMap are provided.
i_Image2BitMap will return a bitmap of the area extracted at a
magnification specified. i_DisposeImage will free the memory
occupied by an image. i_NewImage and i_AddDataToImage can be
used to create an empty image and add compressed data to it.

Here is some example code that prints out an image given its
run-length representation.



              Example 16. Run-length Image Format
  /* begin example: image */
  /* ShowImage creates an image from a bitmap,
    and shows the values of the runs
    in the image. The input bitmap looks like:
  
  11001100
  00001111
  
    The output of ShowImage is:
  0 2 4 6 8 8 -1
  4 8 -1
  */
  static void rlprintline(RlLinePtr s)
  {
    int4 i;
    printf("%lx ", (unsigned long)s);
    for (i = 0; i < s->size - 1; i++)
      printf("%d %d ", s->run[i].up, s->run[i].dn);
    printf("%d", s->run[i].up);
    printf("\n");
  }
  
  static iDAFSError
  rlprintimage(iImagePtr sp)
  {
    int4 i;
  
    for (i = 0; i < sp->height; i++)
      rlprintline(&sp->d.im.rl[i]);
    return DAFSOK;
  }
  
  iDAFSError ShowImage(void)
  {
    char bitmap[2] = {0xcc, 0x0f};
    iDAFSError e;
    iImagePtr im;
    iBox b;
  
    b.x = 0;
    b.y = 0;
    b.w = 8;
    b.h = 2;
    im = i_BitMap2Image(b,1,bitmap);
    if (!im) return MallocFailError;
  
    rlprintimage(im);
    i_DisposeImage(im);
    return DAFSOK;
  }
  /* end example: image */
     
i_CloneEntityImage will return a new image that is the
entity's image content bounded by the entity's enclosing box.

The following example illustrates the use of
i_SetImageCallback, i_GetDataFromImage, and i_AddDataToImage.
i_SetImageCallback is used to write the images of an entity to
separate files. i_GetDataFromImage and i_AddDataToImage are
used to access the CCITT Group 4 compressed data of an image.
This example shows how to write the images into separate files
that are of a type not supported by DAFSlib. The image file
format is simply the bounding box of the image followed by the
CCITT group 4 compressed data.

         Example 17. Storing Images in a Separate File
  /* begin example: separate file and image */
  static iDAFSError  writeproc(char *tag,iImagePtr im)
  {
     int fd = -1;
    iBox b;
    iDAFSError e;
    int2 size;
    char *data;
     int i;
  
    fd = open(tag, O_WRONLY);
    if (fd < 0) {e = WriteError; goto error; }
    b = i_GetImageBounds(im);
    if (write(fd, &b, sizeof(iBox)) != sizeof(iBox)) {
       e = WriteError; goto error;
     }
    for (i = 0; e = i_GetDataFromImage(im, i, &data, &size);
      i++) {
     if (write(fd, data, size) != size) {
          e = WriteError; goto error;
       }
    }
    if (e == NoMoreData) e = DAFSOK;
  error:
    if (fd >= 0) close(fd);
    return e;
  }
  
  #define CHUNKSIZE 4096
  static iDAFSError  readproc(char *tag,iImagePtr *imp)
  {
    iImagePtr im = 0;
    char *data = 0;
    iDAFSError e = DAFSOK;
    iBox b;
    int fd, quan;
  
    fd = open(tag, O_RDONLY);
    if (fd < 0) {e = WriteError; goto error; }
    if (read(fd, &b, sizeof(iBox)) != sizeof(iBox)) {
       e = ReadError; goto error;
     }
    if (!(im = i_NewImage(b.w, b.h, 300, NoComp))) {
       e = MallocFailError; goto error;
     }
    if (!(data = (char *) malloc(CHUNKSIZE))) {
       e = MallocFailError; goto error;
     }
    for (;;) {
     quan = read(fd, data, CHUNKSIZE);
     if (quan <= 0) break;
     if (e = i_AddDataToImage(im, data, quan)) goto error;
    }
     *imp = im;
    im = 0;
  error:
    if (im) i_DisposeImage(im);
    if (data) free(data);
    return e;
  }
  
  /* WriteNReadEntityImage just creates a small all white
     image and attaches it to an entity. It then writes the
     entity, disposes of it and reads it back.
  */
  iDAFSError WriteNReadEntityImage(void)
  {
     iImagePtr im = 0;
    iEntityPtr en = 0;
    char *tag;
    iDAFSError e = DAFSOK;
  
    i_SetImageCallBack(writeproc, readproc);
  
    if (!(im = i_NewImage(32, 32, 300, NoComp))) {
       e = MallocFailError; goto error;
     }
    if (e = i_SetImageTag(im, "image1.dat"))
       /* the tag is used as the file name */
     goto error;
    if (!(en = i_NewEntity(0, 0))) {
       e = MallocFailError; goto error;
     }
       i_SetImage(en,im);
     if (e = i_WriteEntity("test.dr", en, DAFSB)) goto error;
     i_DisposeEntity(en);
    en = 0;
    i_DisposeImage(im);
    im = 0;
  
    if (e = i_ReadEntity("test.dr", &en)) goto error;
    i_SetImageCallBack(0, 0);
  error:
    if (im) i_DisposeImage(im);
    if (en) i_DisposeEntity(en);
    return e;
  }
  /* end example: separate file and image */
     
i_WriteImage writes an image and store it in any of the
supported formats. i_ReadImage will read the image in the file
fileName in any of the supported formats and return a pointer
to it in iImagePtr. The read image can be uncompressed or
CCITT group 3 or 4, and is put into the internal iImage format
when read in. The image data will not be uncompressed until
the data actually needs to be used. i_ReadImage and
i_WriteImage each use the file name extension to determine
what image format the data is or should be stored as. Comp
should be passed as NoComp, CCITT3, or CCITT4 into the write
routine.

Sorting

i_Sort is a function that sorts the children of an entity
according to a comparison function passed to it. The
comparison function follows the convention used by the
standard C function strcmp: it returns 0 if the two entities
are equal, a positive value if the first one is greater, and a
negative value if the second one is greater. Two comparison
functions are supplied. The first function, i_HorzCompare,
sorts the children by their horizontal position in the image.
The second function, i_VertCompare, sorts the children by
their vertical position in the image.

Call Back Mechanism

The programmer can turn on a call back mechanism with the
i_SetCallBack routine. This will make DAFSLib send messages to
the application whenever an entity is created, destroyed, or
changed. i_SetCallBack makes it easy to create an interactive
GUI for DAFSLib and is used by the Illuminator application.

A useful application of callbacks is illustrated in the
example which follows. Let's say that you want to write an OCR
program where the bounding box of a parent entity is
automatically adjusted to be exactly the enclosing rectangle
of all of its children. We accomplish by setting up a callback
for an entity to adjust its parent's bounding box whenever it
moves.

                     Example 18. Callbacks
  /* begin example: callback */
  void adjustParentBounds(iEntityPtr entity)
  {
     iEntityPtr parent, child;
     iBox child_union, b;
     iEntityPtr e;
     child_union.w = 0; /* no bounds are denoted by w == 0 */
     parent = i_GetParentEntity(entity);
     for(child = i_GetChildEntity(parent); child; child =
     i_GetNextEntity(child)) {
       b = i_GetEnclosingBox(child);
       if(b.w) { /* this child has area */
          if(child_union.w == 0) /* and first one */
               child_union = b;
          else /* add to previous children's boxes */
               i_UnionBox(b, child_union, &child_union);
       }
     }
     /* will cause chain reaction, might check if unchanged */
     i_SetBox(parent, child_union);
  }
  
  iDAFSError entityCallback(iMess m, void *user_data,
     iEntityPtr entity)
  {
     switch(m) {
     case mMove:    adjustParentBounds(entity); break;
     /* we could handle more messages if we wanted to */
     }
  }
  
  int main(int argc, char **argv)
  {
     extern void doAllExamples(); /* for internal testing */
     i_InitLib();
     /* this would probably be called right after i_InitLib */
     i_SetCallBack(entityCallback, 0);
     doAllExamples();
     i_ExitLib();
  }
  /* end example: callback */
     


DAFS Low-level I/O Routines

i_ReadToken, i_Read1Token, i_ReadItem, i_ReadItemAlloc, and
i_WriteItem are low level routines that are used by DAFSLib to
read DAFS-BINARY documents.

DAFSLib file I/O is mediated by the ri... series of functions.
They access a table of functions that by default are set up to
refer to the ANSI stream library functions. The riInit
function allows the API programmer to override the File I/O
function table, and supply his own functions. The arguments
are compatible with those stream library for the sake of
convenience.

                    Programmer's Reference
                               
Initialization

  #include "dafs.h"
  
  iDAFSError i_InitLib(void);
  void i_ExitLib(void);
     
i_InitLib must be called to access the library before
calling any other DAFS function. i_ExitLib is called after
the application is through with the library.

Entities

  iEntityPtr i_NewEntity (const iTypeIndex type, iEntityPtr
     parent);
  iEntityPtr i_NewEntityBefore (const iTypeIndex type,
     iEntityPtr sibling);
  iEntityPtr i_NewEntityAfter (const iTypeIndex type,
     iEntityPtr sibling);
  iDAFSError i_CopyEntity (iEntityPtr entity, iEntityPtr
     parent, iEntityPtr *epp);
  void i_DisposeEntity(iEntityPtr entity);
  
i_NewEntity creates a new entity of the type specified and
returns a pointer to it. If parent is not NULL, the entity
is added to the end of the list of subentities of parent. If
parent is NULL, the entity will stand alone. If it can't
allocate enough memory, NULL is returned. The new entity's
image is set to the same image as its parent.

The other new entity routines act similarly.
i_NewEntityBefore inserts the created entity in the slot
previous to the supplied sibling and i_NewEntityAfter
inserts the created entity in the next slot.

i_CopyEntity duplicates the hierarchy rooted at entity and
makes it have parent parent. The new structure is returned
in *epp.

i_DisposeEntity frees all memory associated with the entity,
including all subentities.

  typedef enum {
     DAFSB,
     DAFSU,
     DAFSA
  } iDAFSWriteMode;
  
  iDAFSError i_ReadEntity(const char *file_name, iEntityPtr
     *entity_ptr);
  iDAFSError i_ReadEntityNoIm(const char *file_name,
     iEntityPtr *entity_ptr);
  iDAFSError i_WriteEntity(const char *file_name, const
     iEntityPtr entity, iDAFSWriteMode mode);
     
i_ReadEntity creates a new entity and all subentities by
reading the file file_name. It returns the pointer to the
new entity in entity_ptr. If an error occurs, the file
system error code is returned, otherwise DAFSOK (zero) is
returned. i_ReadEntityNoIm performs the same task except it
does not read the image. This is useful when the entity
contains a large image that is not needed.

i_WriteEntity writes the entity to the file file_name. Mode
specifies whether to write DAFS-ASCII, DAFS-UNICODE or DAFS-
BINARY format with the DAFSB, DAFSA, or DAFSU enumerator.
This function returns an error when the writing was
unsuccessful. Currently, only the DAFSB format is
implemented.

Entity Type

  typedef unsigned char iTypeIndex;
  enum iTypeInt_e {
      iNullT,
      iDocumentT,
      iParagraphT,
      iLineT,
      iWordT,
      iGlyphT,
      iZoneT,
      iRegionT,
      iSpaceT,
      iNewLineT,
      iCommentT,
      iNTypesT /* leave last */
  };
  
  iTypeIndex i_NameToType(const char *type_name);
  const char *i_TypeToName(iTypeIndex type);
     
  void i_SetEntityType(iEntityPtr entity, const char type);
  const iTypeIndex i_GetEntityType(const iEntityPtr entity);
     
The programmer supplies the type_name to i_NameToType and is
returned a valid type index. Type indices are not fixed but
may be modified from one running of DAFSLib to another,
therefore, it is necessary to use this function. An
exception are the predefined types found in enum iTypeInt_e.
Using these predefinied indices is OK.

i_SetEntityType changes entity to type, while
i_GetEntityType returns the type of entity.

Interpreting Child Entities

  typedef enum {
     iAnd,
     iOr
  } iLogicalType;
     
  void i_SetLogicalType(iEntityPtr entity, iLogicalType
     type);
  iLogicalType i_GetLogicalType(iEntityPtr entity);
     
The logical type determines how the subentities should be
interpreted. If the logical type is iAnd, then the
subentities are representations of different items that make
up the parent entity. If logical type is iOr, then the
subentities are all different interpretations of the parent
entity. i_SetLogicalType sets the logical type of entity to
type. i_GetLogicalType returns the logical type of entity.

Entity Transversal

  iEntityPtr i_GetChildEntity(const iEntityPtr entity);
  iEntityPtr i_GetNextEntity(const iEntityPtr entity);
  
  iEntityPtr i_GetLastEntity(const iEntityPtr entity);
  iEntityPtr i_GetPrevEntity(const iEntityPtr entity);
  
  iEntityPtr i_GetParentEntity(const iEntityPtr entity);
  
i_GetChildEntity returns the first subentity of entity, or
NULL if there are no children.

i_GetNextEntity returns the entity after entity or NULL if
it is the last one.

i_GetLastEntity return the last subentity of entity, or NULL
if there are no children.

i_GetPrevEntity returns the entity before entity or NULL if
its the first one.

i_GetParentEntity returns the entity's parent, or NULL if it
has none.

Iterative Searching

  typedef enum {
     ALL,
     POSSET,
     PROPERTY,
     NOCHILD
  } iSearchType;
  
  iEntityPtr i_FindEntityForward(iEntityPtr home, iTypeIndex
     type, iSearchtype method, void *name);
  iEntityPtr i_FindEntityBack(iEntityPtr home, iTypeIndex
     type, iSearchtype method, void *name);
     
i_FindEntityForward and i_FindEntityBack allow searching
through the entity structure for a match to two criteria.
The order traversed is read order. The first criteria is the
entity type index. If the given index is zero all entity
types will be searched. The second search criteria is
supplied by the enumerated variable method:

  ALL: This is a wild card, and will match any entity.
  POSSET: This matches an entity with that has the
     character specified in name at the top of its
     possibility set.
  PROPERTY: This matches entities that have the
     property supplied in name.
  NOCHILD: This matches entities that have no children.
     
For the methods that don't use name, NULL can be supplied.
These functions return the nearest entity in read order from
home that matches both search criteria or zero if no such
entity is found. These functions do not limit themselves to
descendants of home, but can search the entire length of the
document from that point.

Entity Hierarchy Information

  int i_CountChildren(iEntityPtr entity);
  int i_CountPrev(iEntityPtr entity);
  int i_CountNext(iEntityPtr entity);
     
i_CountChildren returns the number of subentities of entity.

i_CountPrev returns the number of siblings previous to
entity. i_CountNext returns the numbers of siblings after
entity.

  i_CountPrev(entity) + i_CountNext(entity) + 1 ==
     i_GetParentEntity(entity) ?
     i_CountChildren(i_GetParentEntity(entity)) : 1
     
Rearranging the Entity Hierarchy

  void i_MoveToFront(iEntityPtr entity, iEntityPtr parent);
  void i_MoveToBack(iEntityPtr entity, iEntityPtr parent);
  void i_MoveBefore(iEntityPtr entity, iEntityPtr
     NextSibling);
  void i_MoveAfter(iEntityPtr entity, iEntityPtr
     PrevSibling);
     
i_MoveToFront moves entity to be the first subentity of
parent.

i_MoveToBack moves entity to be the last subentity of
parent.

i_MoveBefore moves entity so that it is before NextSibling.

i_MoveAfter moves entity so that it is after PrevSibling.

Setting Properties

  typedef enum {
     iPNo,
     iPBool,
     iPLong,
     iPString,
     iPUString,
     iPUserData,
     iPUserPtr
  } iPTag;
  
  typedef void (*iUSerPtrFree)(void *data);
  typedef void (*iUserDataProc) (void *p);
  
  iDAFSError i_SetBoolProp(iEntityPtr entity,
               const char *name,int data);
  iDAFSError i_SetLongProp(iEntityPtr entity,
               const char *name, int4 data);
  iDAFSError i_SetStringProp(iEntityPtr entity,
                const char *name, char *data);
  iDAFSError i_SetUStringProp(iEntityPtr entity,
                const char *name, iUCode *data);
  iDAFSError i_SetUserDataProp(iEntityPtr entity,
                 const char *name,
                 void *data, uint2 size);
  iDAFSError i_SetUserPtrProp(iEntityPtr entity,
                const char *name, void *data,
                iUserPtrFree free_proc);
  
  /* associate property with a type */
  iDAFSError i_SetSessionBoolProp(iTypeIndex itype,
                  const char *name, int data);
  iDAFSError i_SetSessionLongProp(iTypeIndex itype,
                  const char *name, int4 data);
  iDAFSError i_SetSessionStringProp(iTypeIndex itype,
                   const char *name,
                   char *data);
  iDAFSError i_SetSessionUStringProp(iTypeIndex itype,
                    const char *name,
                    iUCode *data);
  iDAFSError i_SetSessionUserDataProp(iTypeIndex itype,
                    const char *name,
                    void *data,uint2 size);
  iDAFSError i_SetSessionUserPtrProp(iTypeIndex itype,
                    const char *name,
                    void *data,
                    iUsrPtrFree freeProc);
  
i_SetXXXProp allocates any memory required for the new
property, adds it to the end of the list of properties
attached to entity. Except for the UserPtr functions all the
functions make a copy of the data. The UserPtr functions
store just a pointer. It is your data and DAFSLib knows
nothing about it,; it will not read it or write it. Other
properties will be saved when the entity is written to a
file. The size parameter is needed for the iPUserData type,
it refers to the length, in bytes, of the contiguous section
of memory pointed to by data. You can supply freeProc to the
UserPtr set functions and DAFSLib will call it when it wants
to free the property. If freeProc is NULL, management of the
data remains in the hands of the application programmer.

i_SetSessionXXXProp attaches the property name to the entity
type described by the index number type. Thereafter,
entities of that type will return the value associated with
the type if there is no value associated with the particular
entity.

Getting Properties

  iDAFSError i_GetBoolProp(const iEntityPtr entity,
               const char *name, int *ip);
  iDAFSError i_GetLongProp(const iEntityPtr entity,
               const char *name, int *ip);
  iDAFSError i_GetStringProp(const iEntityPtr entity,
                const char *name, const char **ip);
  iDAFSError i_GetUStringProp(const iEntityPtr entity,
                const char *name, iUCode **ip);
  iDAFSError i_GetUserDataProp(const iEntityPtr entity,
                 const char *name,
                 void **d, long *size,
                 iUserDataProc convertProc);
  iDAFSError i_GetUserPtrProp(const iEntityPtr entity,
                const char *name, void **d,
                iUserPtrFree *freeProcPtr);
     
The Boolean, Integer, String, and Unicode string data can be
retrieved by i_GetBoolProp, i_GetLongProp, i_GetStringProp,
and i_GetUStringProp respectively.

i_GetUserPtrProp sets *d to the pointer to the user defined
data that was originally supplied, and *freeProcPtr to the
free function that was supplied. For either of these
arguments, feeding a NULL pointer tells DAFSLib not to
bother to return that parameter.

i_GetUserDataProp returns a pointer to a copy that was made
of the original data and stored in the entity. It also
returns the size, the amount of bytes of contiguous memory.
convertProc swaps the data returned if necessary. If the
data is a single byte, or does not need swapping,
convertProc can be NULL.

Disposing of Properties

  void i_DisposeProp(iEntityPtr entity, const char *name);
  void i_DisposeAllProps(const iEntityPtr entity);
  void i_DisposeSessionProp(const iTypeIndex type,
               const char *name);
  void i_DisposeAllSessionProps(void);
     
i_DisposeProp removes the property of type name from the
list and frees any data associated with it, including the
string or the user data if the type is iPString or
iPUserData. If the type is iPUserPtr, it will not be freed.

i_DisposeAllProps disposes of all properties on the
iEntityPtr entity. These functions do not affect properties
set using the i_SetSessionXXXProp functions.

i_DisposeSessionProp will dispose of properties matching
name attached to type.

i_DisposeAllSessionProps disposes of all properties on the
entity type index. These function do not affect properties
set using the i_SetXXXProp functions.

Traversing and Finding Properties

  iPropPtr i_FirstProp(const iEntityPtr entity);
  iPropPtr i_NextProp(const iPropPtr LastProp);
  iPropPtr i_FindProp(const iEntityPtr entity,
            const char *name);
  iPropPtr i_FindPropNoTypes(const iEntityPtr entity,
                const char *name);
     
i_FirstProp returns the first property of an entity, or NULL
is the entity has no properties. This will not return
properties that have been associated with the entity's type
by an i_SetSessionXXXProp function. The order in which
properties are stored is an undocumented implementation
detail.

i_NextProp returns the next property of an entity after
LastProp, or NULL if there aren't anymore.

To find the actual property (not just the value) by name,
use i_FindProp, which will return the property when found or
NULL if not found. It operates in the same way as the get
functions in that if entity does not have the property, it
will return a property of name associated with the type of
entity.

i_FindPropNoTypes is identical to the last function, except
that it ignores session properties.

Property Attributes

  const char *i_GetPropName(const iPropPtr prop);
  iPTag i_GetPropType(const iPropPtr prop);
  long i_GetPropSize(const iPropPtr prop);
  iDAFSError i_ChangePropName(iPropPtr, const char
     *from_name, const char *to_name);
     
The name, type, and size of the property can be retrieved by
calling i_GetPropName, i_GetPropType, and i_GetPropSize
respectively.

You can change the name of a property using
i_ChangePropName. If there already is a property with name
to_name, it is replaced.

Transferring And Copying Properties

  iDAFSError i_MoveProp(iEntityPtr entity, iEntityPtr to_en,
             const char *name);
  iDAFSError i_MoveAllProps(iEntityPtr from_en,
               iEntityPtr to_en);
  iDAFSError i_CopyAllProps(iEntityPtr from_en,
               iEntityPtr to_en);
     
i_MoveProp reassigns the properties of type name from entity
to to_en.

i_MoveAllProps reassigns all the properties attached to
from_en to to_en.

i_CopyAllProps copies each property from from_en to to_en.

Comments

  iDAFSError i_SetComment(iEntityPtr entity,
              const char *comment);
  const char *i_GetComment(const iEntityPtr entity);
     
These routines all get and set information about an entity.
Comments are string data that the programmer wants to keep
with the entity. illumIlluminator displays comments.

Contents

  enum {
     iDNo,
     iDPosSet,
     iDString,
     iDUString,
     iD1PosSet,
     iDUserData,
     iDConf,
     iDInt
  } iDTag;
     
  typedef struct {
     unsigned short conf, chr;
  } iPosSet;
     
  iDTag i_GetEDataType(iEntityPtr entity);
  void i_SetNoData(iEntityPtr entity);
  const char *i_GetText(const iEntityPtr entity);
  iDAFSError i_SetText(iEntityPtr entity,const char *text);
  const iUCode *i_GetUText(const iEntityPtr entity);
  iDAFSError i_SetUText(iEntityPtr entity, const iUCode
     *text);
  short i_GetGlyph(const iEntityPtr entity, short *conf);
  void i_SetGlyph(iEntityPtr entity,short chr, short conf);
  const iPosSet *i_GetPosSet(iEntityPtr entity,int *n);
  void i_SetPosSet(iEntityPtr entity,int n,const iPosSet
     *chrs);
  short i_GetShort(const iEntityPtr entity, short* conf);
  void i_SetShort(iEntityPtr entity, short num, short conf);
  short i_GetConf(const iEntityPtr entity);
  void i_SetConf(iEntityPtr entity, short conf);
     
i_GetEDataType returns what type was actually used.

If the user wants to attach a string to the entity,
i_GetText and i_SetText should be used. i_GetUText and
i_SetUText are the analogues for Unicode strings.

For single characters, the library will use less memory if
i_SetGlyph and i_GetGlyph are used, where the value is a
single Unicode character. This corresponds to the the tag
iD1PosSet.

If there is more than one alternative for the value of the
entity then, i_GetPosSet and i_SetPosSet should be used. For
these routines, n is the number of valid possibilities in
the array pointed to by chrs. The library copies the data
that chrs points to, so that the programmer doesn't have to
allocate the pointer.

The glyph and posset routines can be mixed. For example it
is OK to call i_SetGlyph then call i_GetPosSet. The return
values are undefined if the text or the user routines are
mixed with the others.

i_SetShort is used to store a short with an attached
confidence. The illum programIlluminator expects that
entities of type Space or Newline will store how many spaces
or new lines they represent as shorts.

i_SetConf and i_GetConf are used to get and set the
confidence of any sort of entity contents except possibility
sets.

i_SetNoData can be used to clear all data. It is also
possible to send in a NULL string to i_SetUText or
i_SetText, or a zero-length pos set to i_SetPosSet, or the
NUL character to i_SetGlyph. Supplying a NULL value for
conf, tells DAFSLib not to bother returning that parameter.

Fixed Point Numbers

  typedef int4 iFixedPoint;
  /* fixed point conversions etc. */
  #define i_INT2_TO_FP(x) (((iFixedPoint)(x))<<16)
  #define i_FP_TO_INT2(x) ((int2)((x)>>16))
  #define i_FP_FRAC(x) ((x)&0xFFFFL)
  #define iFP_MAX 0x7fffffff
  #define iFP_MIN (-iFP_MAX-1)
  
  #define iFP_HALF (1L<<15)
  #define iFP_ONE (1L<<16)
  
  #define i_D_TO_FP(x) ((iFixedPoint)((x)*65536.0))
  #define i_FP_TO_D(x) ((double)((x)/65536.0))
  
  /* this macro should probably be a function that does the
     multiply better */
  #define i_MUL_FP(x,y) ( ((x)/256)*((y)/256) )
     
i_INT2_TO_FP converts a short to a fixed point number. The
fractional part of the number will be zero. i_FP_TO_INT2
converts a fixed point number to a short. The fractional
part is lost.

i_FP_FRAC returns the fractional part of the fixed point
number. It is still expressed in parts of 2-16.

iFP_MAX is the largest positive number representable by a
fixed point number. It can be expressed as 32767.9999847.
iFP_MIN is the largest negative number, -32768.9999847.

i_FP_HALF is a convenient macro for the fixed point
representation of 1/2. i_FP_ONE is a convenient macro for
the fixed point representation of 1.

i_D_TO_FP converts a double to a fixed point number. The
range and precision of a fixed point number is less than
that of a double, so this must be used carefully. i_FP_TO_D
converts a fixed point number to a double, which should
never have any trouble representing a fixed point number.

i_MUL_FP is a macro for multiplying fixed point numbers.

Images

Images are stored in a run-length form. Each image is made
up of lines, and each line is made up of a series of
transitions or runs.

  typedef struct {
    short up, dn;
  } transition;
     
In the transition structure, up is the position of the
begining of a run of black pixels, dn is the position of the
pixel after the end of a run of black pixels. So, dn minus
up is the number of black pixels in this run.

  typedef struct {
    int2 size; /* size of this array */
    int2 asize; /* alloc'd size of this array */
    int2 last; /* last looked at transition */
    int2 unused;
    short *code; /* this points to the allocated run array
     */
    transition *run; /* this points to &code[1],
       therefore it shouldn't be freed */
  } * RlLinePtr;
     
The number of transitions in this line is stored in size.
asize is the number of transitions that have been allocated;
usually this is the same as size. last can be used for
caching the last looked at transition for this row. code
points to the allocated space for the transitions. run is
set to &code[1].

  code[0] is always 0.
  run[size - 1].dn is unused.
  run[size - 1].up is always -1.
  run[size - 2].dn is always set to the width of the
     image.
  run[size - 2].up will be set to the width of the
     image if this line ends with a white pixel.
  
  
  typedef struct {
    RlLinePtr rl;
    long pixcnt,trncnt;
    iFixedPoint xmul, ymul;
    int rot;
    int2 x, y, w, h; /* smallest enclosing rectangle */
    struct iImage *next, *prev;
  } iRl;
  
iRl is an array of RlLinePtr. The size of the array is the
height of the image. pixcnt, trncnt, xmul, ymul, rot, x, y,
w, h, next, and prev are not used. These can be used by the
developer.

  enum iType {
    Image, Raw
  };
  
  /* The currently allowed image compression data formats.
     */
  enum iDAFSCompType {
    NoComp,
    CCITT3,
    CCITT4
  };
  
  typdef struct {
    iType type;
    int use;
    int4 id;
    int2 xOff, yOff, height, width, dpi;
    char *tag;
    union {
      iRl im;
      iRaw rw;
    } d;
  } *iImagePtr;
  
  iDAFSError i_CompressImage(iImagePtr image,
                iDAFSCompType comp);
  iDAFSError i_UncompressImage(iImagePtr image);
  
  use contains the number of references to this image.
     The image will not be freed by i_DisposeImage
     until this reaches 0.
  id is used during the writing of an entity.
  xOff, yOff are not used.
  height and width are the image's width and height.
  dpi is the dots per inch of the image.
  tag can be used to store any arbitrary string.
     
These items should not be accessed directly, instead use the
corresponding function.

type determines which element of the union is used. If type
is set to Image then d.im is valid. If type is set to Raw
then d.rw is valid. If the type is Image, then the format
for run-length images explained above is valid. The Raw type
is not documented externally. Use i_UncompressImage to
guarantee that image is in run-length format, and use
i_CompressImage to save space. i_CompressImage is
automatically called before an image is written out to disk.

Using Images With Entities

  void i_SetImage(iEntityPtr entity,iImagePtr image);
  const iImagePtr i_GetImage(const iEntityPtr entity);
  iImagePtr i_CloneEntityImage(const iEntityPtr entity);
     
All appropriate image functions manage a use count
associated with the image so that garbage collection can be
done on the image when appropriate. Some functions return a
new image unassociated with an entity. Images returned from
such functions are given a use count of one, and the
application programmer must call i_DisposeImage to
acknowledge that he is done using it. We will mention which
functions cause an implicit use as we get to them in this
section.

i_SetImage sets the entity to point to image.

i_GetImage will return the image attached to the entity
unless the entity does not have an image, or NULL if there
isn't one. This does not modify the entity or the use count.

i_DisposeImage decrements the use count on an image and
frees up its allocated memory if the use count is zero.

i_CloneEntityImage returns a copy of the portion of the
image that the entity's bounds describe. This image will
have an implicit use count of one so the application
programmer should call i_DisposeImage.

Converting and Manipulating Images

  iImagePtr i_BitMap2Image(iBox bounds, int data_width,
               char *data);
  iDAFSError i_Image2BitMap(iImagePtr image, iBox bounds,
               iFixedPoint mag, long *dataWidth,
               char **data);
  iImagePtr i_RotateImage(iImagePtr image, int rot);
  iDAFSError i_ScaleImage(iImagePtr img, iFixedPoint xmag,
              iFixedPoint ymag);
     
An iImagePtr can point to images of different internal
storage formats. i_BitMap2Image and i_Image2BitMap provide
the ways to get a bitmap from an iImagePtr and to create an
iImagePtr from a bitmap.

In i_Bitmap2Image, the bounds parameter specifies a region
of interest that is to be extracted in the conversion.
data_width is the width of the bitmap in bytes, which is the
amount added to the data pointer to get to the next row of
image. data is a pointer to the bitmap image data.

In i_Image2Bitmap, the bounds parameter is again the region
of interest. If the bounds width is zero, then the whole
image will be converted. i_Image2BitMap will allocate and
return a pointer in data if data initially points to NULL,
otherwise it assumes that the caller is providing enough
space for the bitmap where data points. The amount of space
the caller must supply is given by rounding the width up to
the nearest multiple of 32, multiplying by the height and
adding four more bytes. The number of bytes in each row of
the image is returned in data_width. The returned image is
scaled according to the magnification value, mag. mag is a
fixed point number as discussed earlier. The image created
by i_Bitmap2Image has an implicit use count of 1 and so the
application programmer should call i_DisposeImage.

i_RotateImage will return a pointer to the same image
rotated by the amount specified by rot. Rot can be either 90
or -90 where -90 will make an upright image from one that
would be viewed with the head tipped to the right.

i_ScaleImage scales an image in both the x and y directions.
Xmag and ymag are fixed point numbers, q.v. The image can be
scaled by any amount as long as the biggest coordinate is
less than 32768.

Image Attributes

  void i_SetImageRes(iImagePtr image, int resolution);
  int i_GetImageRes(iImagePtr image);
  int i_GetImageWidth(const iImagePtr image);
  int i_GetImageHeight(const iImagePtr image);
  iBox i_GetImageBounds(const iImagePtr image);
     
The only attributes of an image that are available to the
developer at this time are the resolution (generally in
dpi), width and height of the image. The resolution is
usually included in the header information of most image
formats. If the resolution is unknown it will default to 0
(zero) unless the format is a PDA, in which case the
resolution will be assumed to be 300dpi.

  typedef enum {
     NoComp,
     CCITT3,
     CCITT4
  } iDAFSCompType;
     
  int i_ImageExt(const char *fileName);
  
  typedef iDAFSError (*iImageWriteProc)(char *tag,
                     iImagePtr im);
  typedef iDAFSError (*iImageReadProc)(char *tag,
                     iImagePtr *im);
  
  iDAFSError i_SetImageTag(iImagePtr im, char *tag);
  char *i_GetImageTag(iImagePtr im);
  void i_SetImageCallBack(iImageWriteProc writeProc,
              iImageReadProc readProc);
  void i_GetImageCallBack(iImageWriteProc *writeProcp,
              iImageReadProc *readProcp);
  iDAFSError i_ReadImage(const char *fileName,
              iImagePtr *imagePtr);
  iDAFSError i_WriteImage(const char *fileName,
              iImagePtr image, iDAFSCompType comp);
  void i_DisposeImage(iImagePtr image);
     
  iImagePtr i_NewImage(int2 w, int2 h, int2 res,
             iDAFSCompType comp);
  iDAFSError i_AddData2Image(iImagePtr im, char *data,
                int2 count);
     
i_ImageExt will return 1 if fileName has an extension that
is understood by the i_ReadImage routine. It doesn't open
the file to verify that the contents match the extension.
Images returned by i_ReadImage have an implicit use count of
1. The currently supported file formats are TIFF and PDA,
which have extensions of ".tif" and ".pda" respectively.
More file formats may be supported in the future.

i_NewImage creates an all white image that has a width of w,
a height of h, and a resolution of res. Such an image has an
implicit use count of 1. Creating an all white image is not
very usefull, but with the i_AddData2Image routine you can
add compressed data to that image. comp sets what type of
image data will be added. i_AddData2Image adds count bytes
of compressed image pointed to by data to the image im. The
image must have been created with i_NewImage with comp
either CCITT3 or CCITT4, but not NoComp. If you want to add
uncompressed data use i_BitMap2Image. Since count is a short
you might need to make several calls to i_AddData2Image to
get all the data in. The image will be uncompressed into the
internal format when it is needed.

i_SetImageTag will attach the string tag to the image.
i_GetImageTag will return the tag from the image. This is
used with the image callback routines set with
i_SetImageCallBack. i_SetImageCallBack provides DAFSLib with
two routines that will be called when an image needs to be
read or written from i_WriteEntity and i_ReadEntity.
i_WriteEntity will not write the image into the DAFSB file
if wp is not NULL, instead it will call the routine pointed
to by writeProc. writeProc can then write the image to a
separate file with i_WriteImage or any method the programmer
desires. i_WriteEntity will note this in a marker when
writing out the file. When i_ReadEntity finds such a marker
in a DAFSB file it will call the routine pointed to by
readProc if readProc is not NULL. readProc can call
i_ReadImage or i_NewEntity and i_AddData2Entity to make the
new image from raw compressed data. The image tag is passed
to readProc and writeProc so the tag could be the external
filename.

i_DisposeImage will decrement the use count on an image and
free it if the use count reaches 0.

Entity Bounds

An entity's bounds can be stored in one of three formats,
iABox, iAOrthogonList, and iABoxList. We have already
discussed what a box is, and we will describe exactly what
an orthogon or a box list is later in this section.

  enum iATag {iABox, iAOrthogonList, iABoxList};
  iDAFSError i_GetArea(iEntityPtr entity, void **data, iATag
     *tag);
  iDAFSError i_ForceAreaType(iEntityPtr entity, iATag tag);
     
i_GetArea is used to retrieve the image bounds from an
entity. The format of the bounds are described in tag, and
can be used to cast the data to its structure. The data will
be either iBox, iOrthogonListPtr, or iBoxListPtr. In the
case of the last two, the pointer is directly returned.
Therefore, the structure should be copied before being
changed or disposed. By convention, an iBox with a width of
zero, or a NULL iOrthogonListPtr or iBoxListPtr, or a
iOrthogonListPtr or iBoxListPtr with zero components
represent no bounds.

i_ForceAreaType forces the conversion of the bounds in
entity to the type specified by tag. Converting to a box can
lose information if the bounds are too complex to be
expressed in a box. Although a box list can represent an
area with "holes," orthogon lists do not in our current
implementation.

  iBox i_GetEnclosingBox(const iEntityPtr entity);
  void i_SetBox(iEntityPtr entity, iBox box);
  iDAFSError i_SetBoxList(iEntityPtr, iBoxListPtr box_list);
  iDAFSError i_SetOrthogonList(iEntityPtr entity,
                 iOrthogonListPtr ortho_list);
     
i_GetEnclosingBox will return the smallest box that encloses
the entity's bounds without affecting the stored format. It
is possible that the entity has no bounds, in that case, the
returned box will have a w field equal to zero.

i_SetBox replaces the stored bounds with box. If (box.w ==
0) that by convention means the entity has no bounds.

i_SetBoxList replaces the stored bounds with box_list. An
error is possible, because, as usual, the value is copied,
therefore storage must be allocated and that could fail.

i_SetOrthogonList replaces the stored bounds with
ortho_list. An error is possible, because, the value is
copied, therefore storage must be allocated and that could
fail.

Orthogons

  typedef struct {
     short x, y;
  } iPoint;
     
  typedef struct {
     int n; /* n describes the number of points */
     iPoint *point;
  } *iOrthogonPtr;
     
  iOrthogonPtr i_NewOrthogon (void);
  void i_DisposeOrthogon (iOrthogonPtr orthogon);
  iOrthogonPtr i_CopyOrthogon(iOrthogonPtr orthogon);
  iDAFSError i_UnionOrthogon(const iOrthogonPtr o1,
                const iOrthogonPtr o2,
                iOrthogonListPtr o_list);
  iDAFSError i_IntersectOrthogon(const iOrthogonPtr ortho1,
                  const iOrthogonPtr ortho2,
                  iOrthogonListPtr o_list);
  void i_TranslateOrthogon(iOrthogonPtr ortho,
               short x, short y);
     
Orthogons are polygons in which the angles created by any
three successive points is a right angle.

i_NewOrthogon creates a pointer to a orthogon, and allocates
space for the empty structure. It is possible that there
could be a memory allocation error, in which case the
function will return a NULL pointer.

i_DisposeOrthogon frees all memory used for the orthogon.

i_CopyOrthogon allocates space for a new orthogon copies in
the information in orthogon, and returns the pointer. It is
possible that there could be a memory allocation error, in
which case the function will return a NULL pointer.

i_UnionOrthogon puts the result of the union of ortho1 and
ortho2 into o_list. It has to be returned in an orthogon
list because the resulting shape might be disconnected. The
value is being returned in a pointer, not a pointer to a
pointer, so a valid iOrthogonListPtr, such as might be
created with i_NewOrthogonList, must be passed in.

i_IntersectOrthogon is exactly the same as i_UnionOrthogon,
except it is the intersection of the two orthogons that is
computed.

i_TranslateOrthogon adds the values of x and y to the x and
y coordinates, respectively, of each point in ortho.

  typedef struct {
      int2 num_poly;
      iOrthogonPtr *poly;
  } *iOrthogonListPtr;
  
  iOrthogonListPtr i_NewOrthogonList(void);
  void i_DisposeOrthogonList (iOrthogonListPtr o_list);
  i_OrthogonListPtr i_CopyOrthogonList(iOrthogonListPtr
     o_list);
  
  iDAFSError i_UnionOrthogonList(const iOrthogonListPtr
     o_list1,
                  const iOrthogonListPtr o_list2,
                  iOrthogonListPtr result);
  iDAFSError i_IntersectOrthogonList(const iOrthogonListPtr
     one,
                    const iOrthogonListPtr two,
                    iOrthogonListPtr result);
  iDAFSError i_SubtractOrthogonList(iOrthogonListPtr shapes,
                   iOrthogonListPtr bites,
                   iOrthogonListPtr result);
     
i_NewOrthogonList creates a structure for multiple
orthogons, and return a pointer to it. If there is a memory
allocation failure, a NULL pointer will be returned.

i_DisposeOrthogonList frees the memory used by o_list.

i_CopyOrthogonList allocates memory for an orthogon list,
copies the contents of o_list into it, and returns the
pointer. If there is a memory allocation failure, a NULL
pointer will be returned.

i_UnionOrthogonList puts the result of the union of o_list1
and o_list2 into result. The value is being returned in a
pointer, not a pointer to a pointer, so a valid
iOrthogonListPtr, such as might be created with
i_NewOrthogonList, must be passed in. The results are not
defined if any o_list1 or o_list2 and result are the same
pointer. An error is possible because the operation may
involve memory allocation.

i_IntersectOrthogonList is the same as i_UnionOrthogonList,
except that it is the intersection that is computed. It is
also possible that there will be no intersection so result
will have zero orthogons in it.

i_SubtractOrthogonList is the same as i_UnionOrthogonList,
except that the order is important, the pixels in bites that
overlap with shapes are subtracted.

Boxes

  typedef struct {
     short x, y, w, h;
  } iBox;
  
  void i_IntersectBox(iBox a, iBox b, iBox *result);
  void i_UnionBox(iBox a, iBox b, iBox *result);
     
Because iBox is a simple structure, there is no need for
new, dispose, or copy routines.

i_IntersectBox returns the intersection of a and b in
result. The result of result being a pointer to a or b is
undefined. The result of a or b being an empty box is
undefined. If there is no intersection, result has w field
equal to zero.

i_UnionBox returns the union of a and b in result. The
behaviour of the function is not guaranteed if either a or b
are empty boxes.

  
  typedef struct {
    int2 x1, x2, y1, y2;
  } BOX;
  
  typedef struct {
    int2 size;
    int2 numRects;
    BOX *rects;
    BOX extents;
  } *iBoxListPtr;
  
  iBoxListPtr i_NewBoxList(void);
  void i_DisposeBoxList(iBoxListPtr box_list);
  iBoxListPtr i_CopyBoxList(iBoxListPtr box_list);
  iDAFSError i_UnionBoxList(const iBoxListPtr b_list1,
               const iBoxListPtr b_list2,
               iBoxListPtr result);
  iDAFSError i_IntersectBoxList(const iBoxListPtr b_list1,
                 const iBoxListPtr b_list2,
                 iBoxListPtr result);
  iDAFSError i_SubtractBoxList(const iBoxListPtr source,
                 const iBoxListPtr subtract,
                 iBoxListPtr result);
  i_TranslateBoxList(iBoxListPtr box_list,
                    short x, short y);
  
i_NewBoxList creates a pointer to a list of boxes, and
allocates memory for its structure. If there is a memory
allocation failure, a NULL pointer will be returned.

i_DisposeBoxList frees the memory allocated to the box_list.

i_CopyBoxList allocates the space for a new box list, copies
the contents of box_list into that new space, and then
returns the pointer. If a memory allocation error occurs, a
NULL pointer will be returned.

i_UnionBoxList puts the result of the union of b_list1 and
b_list2 into result. The value is being returned in a
pointer, not a pointer to a pointer, so a valid iBoxListPtr,
such as might be created with i_NewBoxList, must be passed
in. The results are not defined if any b_list1 or b_list2
and result are the same pointer. An error is possible
because the operation may involve memory allocation.

i_IntersectBoxList is the same as i_UnionBoxList, except
that it is the intersection that is computed. It is also
possible that there will be no intersection so result will
have zero boxes in it.

i_SubtractBoxList is the same as i_UnionBoxList, except that
the order is important, the pixels in subtract that overlap
with source are subtracted.

i_TranslateBoxList adds the value of x and y to the x and y
coordinates, respectively, to both points in each of the
boxes in box_list.

Translating Bounds

  /* convertors */
  iDAFSError i_BoxList2OrthogonList(iBoxListPtr b_list,
                   iOrthogonListPtr *o_list);
  iDAFSError i_Box2OrthogonList(iBox box,
                 iOrthogonListPtr *o_list);
  iDAFSError i_Box2BoxList(iBox box, iBoxListPtr *b_list);
  iDAFSError i_OrthogonList2BoxList(iOrthogonListPtr o_list,
                   iBoxListPtr *b_list);
  iDAFSError i_Box2Orthogon(iBox box, iOrthogonPtr
     *orthogon);
  iDAFSError i_Orthogon2BoxList(iOrthogonPtr o,
                 iBoxListPtr *b_list);
  
  /* lossy convertors */
  void i_BoxList2Box(iBoxListPtr reg, iBox *box);
  void i_OrthogonList2Box(iOrthogonListPtr o_list, iBox
     *box);
  void i_Orthogon2Box(iOrthogonPtr orthogon, iBox *box);
     
All of these convertors operate in a straight-forward
manner. All the ones that convert to orthogon lists or box
lists take pointers to pointers because they allocate space.
The convertors to box do not allocate the result and take
only a pointer. Converting between orthogon lists and box
lists generally works, except in the case of areas with
holes in them, which box lists can represent, but which
confuse the current implementation of orthogon lists.

Unicode Characters

This series of functions is modeled on the ordinary isascii
etc., functions except they operate on Unicode characters
instead of ASCII characters.

  int4 isualnum(iUCode uc);
  int4 isualpha(iUCode uc);
  int4 isuascii(iUCode uc);
  int4 isudigit(iUCode uc);
  int4 isulatin1(iUCode uc);
  int4 isulower(iUCode uc);
  int4 isupunct(iUCode uc);
  int4 isuspace(iUCode uc);
  int4 isuupper(iUCode uc);
  int4 isuxdigit(iUCode uc);
     
These functions are used to determine what type of Unicode
character is being tested.

isualnum checks for an alphanumeric character; it is
equivalent to (isualpha(uc) || isudigit(uc)).

isualpha checks for an alphabetic character; it is
equivalent to (isuupper(uc) || isulower(uc)).

isuascii checks for a character in the ASCII character set.

isudigit checks for a digit.

isulatin1 checks for a character in DIS 10646's latin1
character set.

isulower checks for a lower - case character.

ispunct checks for any printable character which is not a
space or an alphanumeric character.

isuspace checks for white - space characters. These are:
space, form-feed ('\f'), newline ('\n'), carriage return
('\r'), horizontal tab ('\t'), and vertical tab ('\v').

isuupper checks for an uppercase letter.

isuxdigit checks for a hexadecimal digits, i.e. one of 0 1 2
3 4 5 6 7 8 9 0 a b c d e f A B C D E F

These functions are used to manipulate Unicode characters.

  iUCode toulower(iUCode uc);
  iUCode touupper(iUCode uc);
     
touupper converts the letter uc to upper case, if possible.

toulower converts the letter uc to lower case, if possible.

Unicode Strings

The next series of functions is modeled on the strcmp, etc.
functions except that they operate on Unicode strings
instead of ASCII.

  iUCode * ucscat(iUCode *ws1, const iUCode *ws2);
  const iUCode * ucschr( const iUCode *ws, iUCode uc);
  int ucscmp(const iUCode *ws1, const iUCode *ws2);
  iUCode * ucscpy(iUCode *ws1,     const iUCode *ws2);
  unsigned long ucscspn(iUCode *ws1, iUCode *ws2);
  unsigned ucslen(const iUCode *ws);
  iUCode * ucsncat(iUCode *ws1, const iUCode *ws2, int2 n);
  int ucsncmp(const iUCode *ws1, const iUCode *ws2, int2 n);
  iUCode * ucsncpy(iUCode *ws1, const iUCode *ws2, int2 n);
  iUCode * ucspbrk(iUCode *ws1, iUCode *ws2);
  const iUCode * ucsrchr(const iUCode *ws, iUCode uc);
  unsigned long ucsspn(iUCode *ws1, iUCode *ws2);
  void ucsreverse(iUCode ucs[]);
  long ucstol(iUCode *ucs, iUCode **ptr, int4 base);
  iUCode * ucsucs(iUCode *ws1, iUCode *ws2);
     
The ucscat function appends the ws1 string to the ws1 string
overwriting the NULL character at the end of ws1, and then
adds a terminating NULL character. The strings may not
overlap, and the ws1 string must have enough space for the
result.

The ucschr function returns a pointer to the first
occurrence of the character uc in the string ws.

The ucscmp function compares the two strings ws1 and ws2. It
returns an integer less than, equal to, or greater than zero
if ws1 is found, respectively, to be less than, to match, or
be greater than ws2.

The ucscpy function copies the string pointed to by ws2
(including the terminating NULL character) to the string
pointed to by ws1. The strings may not overlap, and the
destination string ws1 must be large enough to receive the
copy.

The ucscspn function calculates the length of the initial
segment of ws1 which consists entirely of characters not in
ws2.

The ucslen function calculates the length of the string ws,
not including the terminating NULL character.

The ucsncat function is similar to ucscat, except that only
the first n characters of ws1 are appended to ws1.

The ucsncmp function is similar to ucsncmp, except it only
compares the first n characters of ws1.

The ucsncpy function is similar to ucscpy, except that only
the first n bytes of ws2 are copied.

The ucspbrk function locates the first occurrence in the
string ws1 of any of the characters in the string ws2.

The ucsrchr function returns a pointer to the last
occurrence of the character uc in the string ws.

The ucsspn function calculates the length of the initial
segment of ws which consists entirely of characters in
accept.

The ucsreverse function reverses the Unicode string ucs.

The ucstol fucntion converts a Unicode string to a long
integer. The integer will be in base base. If ptr is not
NULL, ptr will point to the next character of the number
scanned.

The ucsucs function finds the first occurrence of the
substring ws2 in the string ws1. The terminating NULL
characters are not compared.



DAFS and Text

  iDAFSError i_ReadUATable(iUATablePtr *table, char
     *filename);
  void i_DisposeUATable(iUATablePtr table);
  char *i_AllocASCIIText(iEntityPtr en, iUATablePtr table);
  iUCode *i_AllocUText(iEntityPtr en);
  void i_FreeASCIIText(char *text);
  void i_FreeUText(iUCode *text);
     
An iUATablePtr points to a table for converting Unicode
values to a readable ASCII string. For instance, the Unicode
letter  could be replaced by the ASCII string "cent". This
table is read from a file that contains a list of the
Unicode values followed by the ASCII string. A short file
could be:

  0x00B5       "micro"
  0x00B7       "middot"
  0x00AC  "not"
  64      "ampersand"
  0x2126       "ohm"
     
The Unicode value can be in decimal or hex with the '0x'
prefix.

i_AllocASCIIText uses this table to convert the data in an
entity structure to create a equivilent text string.
i_AllocUText converts the entity data into a Unicode string.
Each of these strings should be freed later by the
complementary routine, i_FreeASCIIText or i_FreeUText.

  iDAFSError i_Collapse(iEntityPtr entity, iUCode *nonrec);
     
i_Collapse is used to consolidate the entity structure when
the lower entities like glyphs are no longer needed. It will
attach a string to the specified entity that duplicates the
information from the subentities. The sub entities are then
removed. The second parameter, nonrec,  substitutes in the
supplicd string whereever a nonrec is encountered.

  typedef int4 (iCompareProc)(iEntityPtr *one, iEntityPtr
     *two);
  iDAFSError i_Sort(iEntityPtr entity, iCompareProc
     compare);
  int4 i_VertCompare(iEntityPtr *one, iEntityPtr *two);
  int4 i_HorzCompare(iEntityPtr *one, iEntityPtr *two);
     
i_Sort will sort the children of entity using the compare
function. The compare function supplied should return a
positive value if one should be before two, 0 if they should
be in the same position, and a negative value if two should
be before one. We supply i_VertCompare which will allow the
programmer to sort entities by their vertical position, and
i_HorzCompare to sort entities horizontally.

Override FILE I/O

  typedef void *RafFile; /* really a FILE * */
  
  typedef int4 (*ReadFunc)(char *, int4, int4, RafFile);
  typedef int4 (*WriteFunc)(const char *, int4, int4,
     RafFile);
  typedef RafFile (*OpenFunc)(const char *, const char *);
  typedef int4 (*CloseFunc)(RafFile);
  typedef int4 (*SeekFunc) (RafFile, long, int4);
  typedef long (*TellFunc)(RafFile);
  typedef int4 (*GetcFunc) (RafFile);
  typedef char *(*GetsFunc) (char *, int4, RafFile);
  typedef int4 (*PutsFunc) (char *, RafFile);
  typedef int4 (*EofFunc)(RafFile);
  
  struct IoFunctionTable {
     ReadFunc read;
     WriteFunc write;
     OpenFunc open;
     CloseFunc close;
     SeekFunc seek;
     TellFunc tell;
     GetcFunc getcf;
     GetsFunc getsclone;
     PutsFunc putsf;
     EofFunc eof;
  };
  
  int4 riRead(char *ptr, int4 size, int4 nitems, RafFile);
  int4 riWrite(const char *ptr, int4 size, int4 nitems,
     RafFile);
  RafFile riOpen(const char *filename, const char *type);
  int4 riClose(RafFile);
  int4 riSeek(RafFile handle, long offset, int4 whence);
  long riTell(RafFile handle);
     
DAFSLib file I/O is mediated by the ri... series of
functions. They access a table of functions that by default
are set up to refer to the Stream library functions. The
riInit function allows the API programmer to override the
file I/O function table, and supply his own functions.

The arguments are plug compatible with the Stream library
for the sake of convenience.


Text Matching

  iDAFSError i_MakeARegex(iRegexPtr **rep, const char *str);
  iDAFSError i_MakeURegex(iRegexPtr **rep, const uChar
     *str);
  int i_CheckARegex(const iRegexPtr *re, const char *str,
           int *start, int *end);
  int i_CheckURegex(const iRegexPtr *re, const uChar *str,
           int *start, int *end);
  void i_DisposeRegex(iRegexPtr *re);
  iDAFSError i_ReadRegex(iRegexPtr *rep, char *filename);
  iDAFSError i_ReadRegexFile(iRegexPtr *rep, RafFile file,
                 int *swap);
  iDAFSError i_WriteRegex(const iRegexPtr re, char
     *filename);
  iDAFSError i_WriteRegexFile(const iRegexPtr re, RafFile
     file);
     
i_MakeARegex takes an ASCII string containing a regular
expression and constructs ("compiles") an internal data
structure for i_CheckARegex or i_CheckURegex, which can each
use the structure for substring matching. i_MakeURegex does
the same thing for a regular expression specified by a
Unicode string. Note that the compiled regular expression
stores characters in Unicode, regardless of whether the
input was ASCII or Unicode.

i_DisposeRegex disposes of a (compiled) regular expression.
It frees both the iRegexPtr and the data describing the
state machine.

i_CheckARegex searches an ASCII string for the first
substring matching a (previously compiled) regular
expression. If a match is found, *start is set to the
position of the first character of the match, and *end is
set to one position past the last character of the match.
The return value is non-zero if a match is found, zero if no
match is found. i_CheckURegex does the same thing for a
string in Unicode representation.

i_ReadRegex opens the file specified by filename, and reads
in a regular expression. The file must have been created by
a call to i_WriteRegex or i_WriteRegexFile. The pointer *rep
is set to the read-in regular expression.

i_ReadRegexFile reads in a regular expression from the
already opened file specified by the RafFile object file.
The parameter swap indicates if the file has a different
byte ordering than the computer reading it in. This value
should be passed to any other functions reading from the
same file.

i_WriteRegex creates the file specified by filename, and
writes out a regular expression.

i_WriteRegexFile writes out a regular expression to the
already opened file specified by the RafFile object file.

                Table 3. Regular Expressions
Characte    Usage                Description
   r
<ordinar  <ordinary  An ordinary character (not a
   y     character>  special character. q.v.) is a one
characte             character regex which matches just
   r>                that character.
  '.'        '.'     A '.' (period) is a one-character
                     regex that matches any character
                     except Newline.
  '\'    '\'<specia  A backslash (\) followed by any
              l      special character is a one-
         character>  character regex that matches that
                     special character itself.
'[' ']'  '['<string  A non-empty string of characters
            >']'     enclosed by '[' and ']' (square
                     brackets) is a one-character regex
         '[^'<strin  that matches any character in that
            g>']'    string. If the first character of
                     the string is '^', the regex
                     matches any character except
                     Newline and the remaining
                     characters in the string.
   '-'   '['<ASCII>  The '-' (minus) may be used inside
             '-'     square brackets (above) to indicate
         <ASCII>']'  a range of consecutive ASCII
                     characters. For example [0-9] is
                     equivalent to [0123456789].
  '*'       <reg.    A regex followed by '*' (asterisk)
          exp.>'*'   matches zero or more occurrences of
                     the expression.
  '+'       <reg.    A regex followed by '+' (plus)
          exp.>'+'   matches one or more occurrences of
                     the expression.
  '?'       <reg.    A regex followed by '?' (question
          exp.>'?'   mark) matches zero or one
                     occurrences of the expression.
  '^'     '^'<reg.   A '^' (caret) at the beginning of
            exp.>    an entire regex constrains that
                     expression to match an initial
                     segment of a line. A '^' is special
                     only at beginning of a regex.
  '$'       <reg.    A '$' (currency symbol) at the end
          exp.>'$'   of an entire regex constrains that
                     expression to match a final segment
                     of a line. A '$' is special only at
                     the end of a regex.
   '|'      <reg.    Two regex's separated by '|' match
          exp.>'|'   any string which matches either the
            <reg.    first or second expression.
            exp.>
              Background on Regular Expressions
                              
    Regular expression notation is used to define sets of
     character strings. We say that a string "matches" a
   regular expression if the string is in the set defined
      by the regular expression. The regular expression
     notation supported by DAFSLib is summarized in the
                        table below.
                              
      The characters '.', '*', '[', and '\' are always
       special, except when they appear within square
     brackets. Under certain conditions, the characters
      '+', '?', '-', '^' '$', and '|' are special (see
        below). ']' is special following an open '['.
                              
     A concatenation of regular expressions matches the
   concatenation of the strings matched by each component
                    of the concatenation.
                              
    A regular expression in parentheses matches the same
       strings as the regular expression. The order of
    precedence of operators at the same parenthesis level
       is '[ ]' (character classes), then '*' '+' '?'
          (closures), then concatenation, then '|'
                  (alternation)and Newline.
                              
       For more detailed information, see "Compilers:
    Principles, Techniques, and Tools" by Aho, Sethi, and
    Ullman or UNIX documentation for the "grep" command.
                              
                         Call Backs
                              
                       typedef enum {
                            mUpdate,
                           mDestroy,
                            mCreate,
                          mWillUpdate,
                           mWillMove,
                             mMove,
                            mDestImg
                          } iMess;
                              
     typedef iDAFSError (*iCallBackProc)(iMess message,
                                void *userData,
                              iEntityPtr entity);
       void i_SetCallBack(iCallBackProc callBack, void
                         *userData);
      void i_GetCallBack(iCallBackProc *callBack, void
                        **userData);
 iDAFSError i_DoCallBack(iMess message, iEntityPtr entity);
                              
i_SetCallBack is used to set a global call back routine that
    the library will call any time an entity is created,
  changed, or destroyed. With i_SetCallBack, the programmer
   can create sophisticated user interfaces that use this
 library. The user-supplied call back routine will be passed
    three arguments. The first argument is message which
indicates what happened or is about to happen to the entity.
 The second argument is userData which is the same userData
   that was passed to i_SetCallBack. The third argument is
entity. When the entity is changed, the mUpdate message will
   be sent after the entity has been updated. This should
  trigger the redrawing of all views that show the entity.
   After an entity is created, the mCreate message will be
sent, and before an entity is destroyed the mDestroy message
 will be sent. The i_DisposeEntity routine will recursively
 destroy all children of the entity, so an mDestroy message
   will be generated for each child. To remove a call back
 routine pass NULL in callBack. DAFSLib is reentrant so any
  routine can be called from any of the call backs. If the
 call back routine isn't reentrant, it must remove the call
     back routine before it calls any DAFSLib routines.
                              
       i_DoCallBack calls the call back routine set by
  i_SetCallBack. The message and entity are passed along to
 the routine along with the user data set by i_SetCallBack.
                              
 i_GetCallBack returns the current values for the call back
                   and for the user data.
                              
                        Byte Swapping
                              
 To make DAFS files transportable between computers that use
different byte orders when storing files, these routines are
  included to enable swapping bytes in common data formats.
 DAFS functions such as i_ReadEntity and i_ReadImage already
  use these function to transparently swap data bytes when
  needed. The exception to this is for the data used in the
 property type iPUserData, and is explained in the property
    section. The byte swapping routines are listed below,
     followed by the type of data on which they operate.
                              
    short i_SwapS(short s);                         short
       long i_SwapL(long l);                      long
  void i_SwapLArray(long *array, int num);        array of
                            longs
 void i_SwapPos(iPosSetPtr p);                   iPosSetPtr
                              
   Standard Functions for converting iPUserData are listed
 below, followed by the type of data on which they operate.
                              
      void i_SwapUShort(void *s);                short
       void i_SwapUlong(void *l);                 long
       void i_SwapUBox(void *b);                  iBox
                              
                       Error Messages
                              
                           enum {
                          DAFSOK = 0,
                        FirstError =-99,
                       MallocFailError ,
                     NonExistentFileError,
                     WriteLockedFileError,
                      ProtectedFileError,
                         DiskFullError,
                      DAFSWriteModeError,
                         WrongPropType,
                          WriteError,
                           ReadError,
                          BadFileType,
                           UserStop,
                        NonExistentProp,
                     UnknownDAFSError    ,
                         NotEnoughData,
                        NotInitialized,
                          BadRotError,
                        CantRotateError,
                         BadScaleError,
                       WrongImageFormat,
                         WrongFileType,
                        AlreadyOrEntity,
                         DiffImageType,
                        BadBoundsError,
                        BadEntityPassed,
                        CantLimitError,
                     OutOfBoundsError    ,
                           LastError
                        } iDAFSError;
          const char *i_GetError(iDAFSError error);
                              
   Descriptions of each error are returned by i_GetError.
 Future implementations may change these error codes. DAFSOK
                    will always be zero.
                              
                DAFS-BINARY Utility Routines
                              
  These routines are only useful for reading a DAFS-BINARY
 document. Most developers will not need these routines but
      instead will use i_ReadEntity and i_WriteEntity.
                              
                typedef unsigned char iToken;
                              
       iDAFSError i_Read1Token(RafFile fo, iToken *t);
  iDAFSError i_ReadToken(RafFile fo, iToken *t, int *swap);
   iDAFSError i_ReadSize(RafFile fo, iToken t, long *size,
                                    int *swap);
  iDAFSError i_ReadItem(RafFile fo, iToken t, unit2 space,
                       uint2 *size, void *data, int *swap);
      iDAFSError i_ReadItemAlloc(RafFile fo, iToken t,
                               uint2 *size, void **data,
                                      int *swap);
   iDAFSError i_ReadItemAllocH(RafFile fo, iToken t, long
                           *size,
                              void huge **data, int *swap);
  iDAFSError i_WriteItem(RafFile fo, iToken t, uint2 size,
                                 const void *data);
  iDAFSError i_WriteItemH(RafFile fo, iToken t, long size,
                               const void huge *data);
                   void smfree(void *ptr);
                void smfreeH(void huge *ptr);
                              
 i_Read1Token reads one token from the stream fo and returns
                          it in t.
                              
   i_ReadToken eats up tokens and associated data until it
 finds one it understands, it then returns that token in t.
If it encounters a ByteSwap token, it changes the swap value
    to 1 if the data should be swapped or 0 if not, then
continues looking for a token it understands. If no ByteSwap
  token is found swap is unchanged. Both of these routines
  will return the token with the size of SOD still OR'd in.
                              
i_ReadSize will return in size the amount data that follows.
After this routine, fread must be called to read in the data
              before another token can be read.
                              
     i_ReadItem and i_ReadItemAlloc can be called after
  i_Read1Token or i_ReadToken to read the SOD and then read
 the data. i_ReadItemAlloc allocates space for the incoming
data and returns a point to it in data. i_ReadItem reads the
 data into data unless data is NULL in which case it fseek's
                       over the data.
                              
  i_WriteItem writes a token, SOD, and the data to the file
                             fo.
                              
  The data allocated by i_ReadItemAlloc must be freed using
                  the the smfree function.
                              
i_ReadItemAllocH, i_WriteItemAllocH, and smfreeH are special
   functions for 16-bit Windows compatibility. Since this
   release does not officially support Windows, we are not
                      documenting them.
                              
                Index of Functions and Macros
                              
i_AddDataToImage, 27, 29,       i_FP_FRAC, 45
 30                             i_FP_TO_D, 45, 46
i_AllocASCIIText, 59            i_FP_TO_INT2, 45
i_AllocUText, 59                i_FreeASCIIText, 59
i_BitMap2Image, 27, 28, 48,     i_FreeUText, 59
 50                             i_GetArea, 24, 26, 51, 52
i_BorrowEntity, 1               i_GetBoolProp, 41
i_Box2BoxList, 25, 55           i_GetCallBack, 63, 64
i_Box2Orthogon, 25, 55          i_GetChildEntity, 19, 20,
i_Box2OrthogonList, 25, 55       33, 37
i_BoxList2Box, 25, 55           i_GetComment, 14, 18, 43
i_BoxList2OrthogonList, 25,     i_GetConf, 44, 45
 55                             i_GetDataFromImage, 29, 30
i_ChangePropName, 43            i_GetEDataType, 13, 21, 44
i_CheckARegex, 61               i_GetEnclosingBox, 17, 24,
i_CheckURegex, 61                33, 51
i_CloneEntityImage, 29, 48      i_GetEntityType, 24, 36, 37
i_CompressImage, 47, 48         i_GetError, 65
i_CopyAllProps, 19, 43          i_GetGlyph, 10, 13, 21, 44
i_CopyBoxList, 25, 54, 55       i_GetImage, 17, 24, 27, 30,
i_CopyEntity, 1, 9, 22, 35       48, 49, 50
i_CopyOrthogon, 25, 52, 53      i_GetImageBounds, 27, 30,
i_CopyOrthogonList, 25, 53       49
i_CountChildren, 38, 39         i_GetImageCallBack, 50
i_CountNext, 38, 39             i_GetImageHeight, 27, 49
i_CountPrev, 38, 39             i_GetImageRes, 49
i_D_TO_FP, 45                   i_GetImageTag, 24, 50
i_DisposeAllProps, 42           i_GetImageWidth, 27, 49
i_DisposeAllSessionProps,       i_GetLastEntity, 19, 37
 42                             i_GetLogicalType, 22, 37
i_DisposeArea, 25               i_GetLongProp, 41
i_DisposeBoxList, 25, 54,       i_GetNextEntity, 19, 20,
 55                              33, 37
i_DisposeEntity, 9, 31, 35,     i_GetParentEntity, 19, 20,
 64                              33, 37, 39
i_DisposeImage, 27, 28, 31,     i_GetPosSet, 13, 44
 47, 48, 49, 50, 51             i_GetPrevEntity, 19, 37
i_DisposeOrthogon, 25, 52,      i_GetPropName, 43
 53                             i_GetPropSize, 43
i_DisposeOrthogonList, 25,      i_GetPropType, 43
 53                             i_GetShort, 44
i_DisposeProp, 42               i_GetStringProp, 19, 41
i_DisposeRegex, 61              i_GetText, 6, 10, 12, 21,
i_DisposeSessionProp, 42         44
i_DisposeUATable, 59            i_GetUserDataProp, 16, 17,
i_DoCallBack, 63, 64             41
i_ExitLib, 33, 35               i_GetUserPtrProp, 41
i_FindEntityBack, 20, 38        i_GetUStringProp, 41
i_FindEntityForward, 20,        i_GetUText, 11, 44
 21, 38                         i_HorzCompare, 32, 59
i_FindProp, 42                  i_Image2BitMap, 17, 27, 48,
i_FindPropNoTypes, 42            49
i_FirstProp, 42, 43             i_ImageExt, 27, 50
i_ForceAreaType, 25, 51         i_InitLib, 33, 35
i_INT2_TO_FP, 45                i_SetEntityType, 19, 24,
i_IntersectBox, 25, 54, 55       36, 37
i_IntersectBoxList, 25, 54,     i_SetGlyph, 10, 11, 13, 22,
 55                              44, 45
i_IntersectOrthogon, 25,        i_SetImage, 23, 24, 29, 31,
 52, 53                          48, 49, 50
i_IntersectOrthogonList, 53     i_SetImageCallBack, 23, 31,
i_MakeARegex, 61                 50
i_MakeURegex, 61                i_SetImageRes, 49
i_MoveAfter, 21, 39             i_SetImageTag, 24, 31, 50
i_MoveAllProps, 19, 43          i_SetLogicalType, 22, 37
i_MoveBefore, 21, 39            i_SetLongProp, 18, 40
i_MoveProp, 43                  i_SetNoData, 11, 44, 45
i_MoveToBack, 21, 22, 39        i_SetOrthogonList, 24, 51,
i_MoveToFront, 21, 39            52
i_MUL_FP, 45, 46                i_SetPosSet, 13, 44, 45
i_NameToType, 24, 36            i_SetSessionBoolProp, 40
i_NewBoxList, 25, 54, 55        i_SetSessionLongProp, 40
i_NewEntity, 9, 10, 22, 31,     i_SetSessionStringProp, 19,
 35, 51                          40
i_NewEntityAfter, 35            i_SetSessionUserDataProp,
i_NewEntityBefore, 35            40
i_NewImage, 27, 30, 31, 50      i_SetSessionUserPtrProp, 40
i_NewOrthogon, 25, 52, 53       i_SetSessionUStringProp, 40
i_NewOrthogonList, 25, 53       i_SetShort, 44, 45
i_NextProp, 42, 43              i_SetStringProp, 19, 40
i_Orthogon2Box, 55              i_SetText, 10, 11, 12, 13,
i_Orthogon2BoxList, 55           44, 45
i_OrthogonList2Box, 25, 55      i_SetUserDataProp, 17, 40
i_OrthogonList2BoxList, 25,     i_SetUserPtrProp, 40
 55                             i_SetUStringProp, 40
i_Read1Token, 34, 66            i_SetUText, 11, 44, 45
i_ReadEntity, 23, 31, 36,       i_Sort, 32, 59
 51, 64, 65                     i_SubtractBoxList, 25, 54,
i_ReadEntityNoIm, 36             55
i_ReadImage, 27, 31, 50,        i_SubtractOrthogonList, 25,
 51, 64                          53, 54
i_ReadItem, 34, 66              i_SwapL, 16, 64
i_ReadItemAlloc, 34, 66         i_SwapLArray, 64
i_ReadItemAllocH, 66            i_SwapPos, 64
i_ReadRegex, 61                 i_SwapS, 16, 64
i_ReadRegexFile, 61             i_SwapUBox, 64
i_ReadSize, 66                  i_SwapUShort, 64
i_ReadToken, 34, 66             i_TranslateBoxList, 54, 55
i_ReadUATable, 59               i_TranslateOrthogon, 52, 53
i_RotateImage, 48, 49           i_TypeToName, 24, 36
i_ScaleImage, 48, 49            i_UncompressImage, 6, 47,
i_SetBoolProp, 40                48
i_SetBox, 24, 33, 51, 52        i_UnionBox, 25, 33, 54, 55
i_SetBoxList, 24, 51, 52        i_UnionBoxList, 25, 54, 55
i_SetCallBack, 32, 33, 63,      i_UnionOrthogon, 25, 52,
 64                              53, 54
i_SetComment, 14, 18, 43        i_UnionOrthogonList, 25,
i_SetConf, 44, 45                53, 54
i_VertCompare, 32, 59
i_WriteEntity, 23, 31, 36,
 51, 65
i_WriteImage, 27, 31, 50,
 51
i_WriteItem, 34, 66
i_WriteItemH, 66
i_WriteRegex, 61
i_WriteRegexFile, 61
iFP_HALF, 45
iFP_MAX, 45
iFP_MIN, 45
iFP_ONE, 45
isualnum, 56
isualpha, 56
isuascii, 56
isudigit, 56
isulatin1, 56
isulower, 56
isupunct, 56
isuspace, 56
isuupper, 56
isuxdigit, 56
riClose, 60
riOpen, 60
riRead, 60
riSeek, 60
riTell, 60
riWrite, 60
toulower, 57
touupper, 57
ucscat, 57, 58
ucschr, 57
ucscmp, 57
ucscpy, 57, 58
ucscspn, 57
ucslen, 57, 58
ucsncat, 57, 58
ucsncmp, 57, 58
ucsncpy, 57, 58
ucspbrk, 57, 58
ucsrchr, 57, 58
ucsreverse, 57, 58
ucsspn, 57, 58
ucstol, 57, 58
ucsucs, 57, 58
                              
   Appendix A: Trademarks, Copyrights and Acknowledgments
                              
The use of the term Unicode Standard in this document refers
to The Unicode Standard Worldwide Character Encoding,
Version 1.0, The Unicode Consortium, Addison-Wesley, 1991.
It also incorporates the Unicode 1.0.1 Addendum.

PDA (Processed Document Architecture) is a trademark of
Calera Recognition Systems, Inc.

The use of the term TIFF in this document refers to the
Tagged Image File Format, invented by Aldus.

The distribution for DAFSLib includes acknowledgments for
source code.


