

                   RRR Sorority Division presents 


                               PIMPPA 
 
                        v0.5.2 / 29.May.2001


                    http://pimppa.sourceforge.net/


1. Legal junk
-------------
See file COPYING


2. What is this anyway
----------------------
Imagine the unfortunate situation where much more stuff
(and even more crap, spam, etc) comes from the Net than you are 
bothered to handle manually. Shall we say, several hundred 
files daily, totalling to over hundreds of thousands of files 
in the long run. Direct net connection and access to an
adequate news server easily provides such conditions.

PIMPPA is designed to help in such a case, to minimize the required
human interaction, striving for complete "hands free" operation,
where you could once a week (for example) drop by and check out
the already filtered and polished catch. The system is suited 
for all types of files, but it was first designed with pics in 
mind. :)

PIMPPA is able to automatically fetch and process files from
newsgroups and FTP sites, while performing the necessary
decoding, duplicate discarding and further file processing.

For example, files downloaded by pimppa tools can be automatically 
deposited to appropriate directories based on their filenames.

PIMPPA also provides nifty command line interfaces for 
file access and backup operations. In addition, there's a
Gnome GUI.


3. Requirements
---------------

Must have:

 - Linux 2.x.x
 - MySQL (v3.22.30 will do). 
   You can get them free for noncommercial use from "http://www.tcx.se/"
 - Suck (v4.2.2 will do) for downloading newsgroups. It should be 
   available at "ftp://metalab.unc.edu/pub/Linux/system/news/transport/"
 - Uudeview (v0.5.13 will do) for decoding news. It should be 
   available at "ftp://metalab.unc.edu/pub/Linux/utils/text/"

Optional:

 - fget (v1.1 will do) for ftp leeching. From
   "http://www-dev.cso.uiuc.edu/fget/". Remember that it must
   be patched for pimppa (patch included in "patches/")!
 - GNOME for GUI. See "http://www.gnome.org/"

PIMPPA depends on so much external stuff because there is no reason to
reinvent the wheel. Its a better idea to use the best applications
available for specific tasks. When the external programs get better,
PIMPPA gets better. If completely new programs surface (performing
similar tasks) they could be easily integrated into the system.


4. Install & Upgrade
--------------------

See file INSTALL. "Third floor, second door on the left". 
Isn't bureacracy such a nice thing?


4.1 Uninstall
-------------
There's a script "uninstall_db" in "extras/" to destroy all
pimppa related stuff from MySQL databases. Use that, and
then remove the pimppa installation, e.g. use "make uninstall" 
in the source distribution main dir or "rpm -e pimppa", 
depending on what you're using.


5. How to use
-------------
Here's some examples and hints.


5.1 Getting started on newsgroup sucking
----------------------------------------
(If you dislike hot command line action, you might
want to use the GUI "bowser" instead of following these
instructions. You'll still have to do things in the
same order: Settings -> Preferences -> Add some fileareas 
-> Add some newsgroups. After that you can probably leech.)

To leech stuff from the news, you must first create 
some destination (incoming!) areas. Use the command:

shell> pnewarea -i -n <area_name> -p <directory_path>

As you now have a couple of areas, assign some newsgroups
to them. Any number of newsgroups may be assigned to a single
area.

shell> pnewgrp -n <newsgroup_name> -d <area_name>

Ok! Then its time to leech! 

shell> pleech -v -H <your.news.server>

If it succeeds, you should have several files decoded to
the destination fileareas. Most of them will be crap, 
but lets pretend that there's some famous Goatlord-stuff 
(pictures) called "blah*.jpg" there you wish to keep.

There's not yet any location for Goatlord's material, 
so you would want to create a new filearea for them,
a "non-incoming" area which is meant only for quality stuff.

shell> pnewarea -n goatlord -p /stuff/goatlord/

Now you have a new filearea called "goatlord". Enter the
directory where you got the "blah*" files and execute
"pmv blah* goatlord". Files will be moved, and in future all
files fetched by "pleech" matching a special assign pattern 
created for "blah*" will be automatically moved to the 
"goatlord" filearea, and not to the newsgroup's default
destination.

If rest of the files you got were junk, you can delete them
(or to be really sure that you'll never see them again in
the future, pmv them to *DOOM*). See chapter 7 for details. 

Now it might be a good idea to make "crond" execute 
"pleech -q -H <your.news.server>" nightly. :)


5.2 Files from other sources
----------------------------
You can just normally "mv" the files to some directory 
assigned as a PIMPPA-filearea. Then use "padopt" to add them 
to the PIMPPA database. Later, those files are just as any which
were sucked from the news. Note that adoption will do some
duplicate checking and may discard files.

But suppose you download a lot of unsorted crap to a single
directory, which is not a PIMPPA filearea. Then you can use

shell> padopt -m <area_id>

where <area_id> is the default destination location for
these files. WARNING: all unsuitable files from the
current directory will be deleted (dupes, invalidly named, etc)
and the rest is moved according to knowledge, or to default.


5.3 Viewing & backupping files
------------------------------
See chapter "Scripts" for quick descriptions of some
viewing interfaces.

When your harddisk starts to burst with all the stuff you have
downloaded, it might be time to store them elsewhere. See
chapter "Utilities" for util called "pbackup". 


5.4 Manual duplicate checking
-----------------------------
While surfing the web, you encounter some tasty files, 
named "blood_*.jpg". Unfortunately, you can't remember if
you already have them. (It might be hard sometimes). 
Easy, util "pf" is just for that purpose. Use

shell> pf -p blood%

and you get the result rather quick even if you had over
100000 files on your PIMPPA system.

If the files are online, you could view them just as speedily
by using "pv_name", e.g.

shell> pv_name blood%


5.5 Example for a true fanatic
------------------------------
In the wide world there might be just a few series of files you want,
but you just don't bother to personally seek them out. Well,
just create the assign patterns for the files you want (for
example by entering one file of the wanted series in the database 
and running "passign" or using "bowser") and then always run 
"pleech" with switch "-n". All files not matching your 
existing patterns will be quietly discarded. One day you may be 
lucky and hit your pot of gold. Or perhaps the precious ones drip 
to your harddisk only slowly. ;)


5.6 Changing preferences
------------------------
PIMPPA stores most of its preferences in table p_misc,
reminiscent of infamous registry on windoze. Utilities
will look for (key,value) pairs from the database and 
use them if found. Otherwise compile time defaults are
used. Command line values override both.

You can set preferences using "pcfg -k <key> -d <value>"
and view them by just "pcfg". You can also use the GUI,
see the next chapter.

Keys suitable for modification are:

CFG_BOWSERSINCE

    Value: Number. How many days since does bowser display
    the files on startup. Default: 1

CFG_FILETYPES

    Value: String. Extended RegExp pattern of allowed filetype 
    extensions. When downloading newsgroup articles, we try to 
    match any of these extensions from the Subject: line, and 
    find out the filename that way. Example:
    
    \.jpg|\.png|\.gif|\.mpg|\.avi|\.asf|\.rar|\.r[0-9][0-9]|\.s[0-9][0-9]

CFG_MINASSIGNNAMELEN

    Value: Number. Filename must have this many characters
    before the extension for the assign pattern to be created/noted 
    at all. Short patterns tend to go wrong often. Default: 4
    (means: don't create assign pattern for e.g. file a10.jpg)
    
CFG_NEEDNUMBERS

    Value: Number. Require this many numbers in all filenames,
    before the extension. (example: blah123.r01 == 3 numbers)
    Default: 0. My suggestion: 2

CFG_NEEDOTHERS

    Value: Number. Require this many non-number characters in all
    filenames, before the extension. (example: blah123.r01 == 4 others)
    Default: 0. My suggestion: 2

CFG_NEWSWAITAFTER

    Value: Number. After leeching this many articles, wait 
    CFG_NEWSWAITSECS (to allow the server to catch its breath). 
    Default: 10

CFG_NEWSWAITSECS
    
    Value: Number. After leeching CFG_NEWSWAITAFTER messages,
    wait this many seconds. Default: 2
    
CFG_NNTPSERVER

    Value: String. DNS name of the NNTP news server you're using 
    (like nntp.asdfg.com) 
    
CFG_NNTPUSER
    
    Value: String. Your user name on the news server (can be empty)

CFG_NNTPPASSWORD

    Value: String. Your password on the news server (can be empty)

CFG_NOSPACE
    
    Value: Boolean. Do NOT allow filenames containing spaces. 
    1 == TRUE, 0 == FALSE. Default: 0. My suggestion: 1.

CFG_STRICTMD5
    
    Value: Boolean. In case of MD5 collision, delete the newer file. 
    1 == TRUE, 0 ==FALSE. Default: 1.

CFG_TMPDIR
    
    Value: String. The directory where pimppa keeps its temporary
    files, like uucoded material downloaded from the news. 
    This directory should usually have atleast several hundred 
    megabytes of free space. Default: "/tmp/"

CFG_TOLOWERCASE
 
    Value: Boolean. Convert incoming filenames to lowercase. 
    1 == TRUE, 0 == FALSE. Default: 1

CFG_VIEWER

    Value: String. Use this viewer executable to view files in
    bowser. It must view the files in the current directory.
    Default: gqview
    
    
6. GUI
------
PIMPPA has a Graphical User Interface called "bowser", 
which lets you download newsgroups and perform various 
operations on your fileareas, files, assign patterns, 
file types, newsgroups and other settings. The GUI usage 
should be self-explanatory. Most operations (like View/Delete/Move) 
are accessed by pressing the right mouse button on a list window.

NOTES:

1) GUI uses "gqview" to view files. It can be changed
   from Settings/Preferences/Miscellaneous. 

2) Warning: "View" and "Delete" operations automatically 
   affect all selected items! They do not ask confirmations.

Most of the really powerful operations of PIMPPA are 
done using plain shell utilities and scripts - without 
fancy user interfaces. :) So after you get bored with 
the GUI, take a closer look at the commandline stuff.


6.1. Setting preferences from GUI
----------------------------------
All fields in prefs reflect their respective SQL tables
and table contents.

With a little thinking and reading this file you should 
be able to get 'em right. Remember that you have to
restart the GUI for some of the "Miscellaneous" preferences
to take effect.

NOTES:

1) When adding fileareas with the GUI, remember
   to include "/" to the end of the directory path.

2) Remember to set all areas which newsgroups are directed
   to as INCOMING.


7. Utilities
------------

    padddir (obsolete at the moment, see Note)
    -------
    This adds the contents of the current directory to filearea 0.
    It can be used to prevent the stuff there from bothering you
    ever again (they will be skipped if re-encountered in future).

    Usage:  padddir [-v]             
    
            -v              Verbose execution

    Note: Perhaps a better way to do this would be to move
    the files to some 'dump' filearea, add them there with
    'padopt' and then kill them with 'prm'. That would produce 
    a couple of assign patterns instead of keeping all the filenames 
    in the database. Result is more general: all similarly named 
    files will be deleted in the future. You can also move
    unwanted files to *DOOM*. This keeps the MD5 checksums of
    the files as well (preventing renamed versions of the
    same files surfacing again).


    padopt
    ------
    This goes through all the PIMPPA fileareas and adds the
    files that are found in the dirs, but not in the database, 
    to the database.
    
    Usage:  padopt [-v]
   
            -a <pattern>    SQL area name pattern to operate on, def=all
            -m <area_id>    Move files from current directory (unless
                            dest. for file is known, move to <area_id>)
            -v              Verbose execution
    
    -m switch moves files matching assign patterns from current 
    directory (which shouldn't be any filearea!) to fileareas and
    the rest to area_id. WARNING: Unsuitable files and duplicates 
    are deleted!

    Note: padopt updates assign patterns.


    passign
    -------
    This browses through all your files and automatically recreates
    the "assign patterns" for them. See section 9.1 for explanation.

    Usage:  passign [-c]

            -c              Delete all previous patterns. Does not
                            affect patterns with destinations 0 or -1.
            -0              Delete all assign patterns with target area 0.
            -n              Delete all assign patterns with target area -1.

    Note: Assign patterns are not created or used for 
    contents of AREA_INCOMING or AREA_NOASSIGN areas.


    pbackup
    -------
    The backup util.

    Files which have "file_backup==0" in database [or the value specified
    with "-r" on the commandline] will be included in the backup.
    Backuped files will be given a new file_backup ID [or the one 
    specified with "-r"] ... 

    Files will be included until media size is reached. Media
    size can be specified on the commandline (in Megabytes).
    In normal operation the next backup will start from the
    area where the last backup operation ended.

    A shadow directory of the backuped files is created
    to a specified location ("-p path") along with a filelist.
    Default is "/tmp/test/". You can then create an ISO-image 
    out of the created directory structure, stream it to 
    tape or whatever.
    
    e.g. after pbackup
    shell> cd /tmp/test
    shell> mkisofs -J -a -d -D -L -R -r -f -o /tmp/cd.img -V JUNK03 .
    shell> cdrecord -data /tmp/cd.img
    
    And to get rid of the backuped files from HD, use

    shell> pclean -i <backup_id>
   
    Note1: Without "-n", this utility tries to fill the backup media 
    to the brim. In any case, it can't know which files should 
    go together. For example, disk archives belonging to the 
    same product may end up on different backups.

    Note2: After using this utility, its suggested to do atleast 
    "du -L" in the destination directory to see that everything
    went ok.
   
    Note3: Contents of INCOMING areas will not be backupped.
    
    Usage:  pbackup [options]

            -n             Be nice, stop at first file not fitting on media 
            -p <path>      Path to make the shadow directory into
            -r <ID>        Specify an explicit backup ID number
            -R             Randomize area order
            -c <area_id>   Skip this file area, can be entered many times
            -s <size>      Media size in megabytes (1MB==1024*1024 bytes)


    pcfg
    ----
    A simple command line tool to configure pimppa, that is,
    to modify p_misc table. You can also use "bowser" or
    "mysql" directly.
    
    Usage:  pcfg [options]

            -k <key>       MUST, the configuration key
            -d <data>      MUST, the value associated with this key
    
    Without options it prints the current setup. See 5.6 
    for known keys.


    pchkfn
    ------
    Checks if a filename can be accepted to system (including
    rules, duplicate & assign pattern check).
    
    Usage:  pchkfn [options] <filename>

            -a <area_id>    An INCOMING area of which's destination
                            regexp patterns are used
            -s              Operate as kill program for suck
            -n              Nazi mode, returns nonzero if no positive 
                            destination area for file is already known

    NOTE: -s is obsolete. PIMPPA won't call "pchkfn -s" anymore,
    though you could use it if you run suck manually by hand with
    your own killfile, with PROGRAM=pchkfn -s.

    Returns 0 if the filename is ok.


    pclean
    ------
    Performs various delete operations. Use with caution!

    Usage:  pclean [options] 
    
            -a <area_id>    Purge duplicates from <area_id>
            -d              Delete files found in db from the current dir
            -f              Delete files which failed integrity check 
            -i <backup_id>  Delete files associated with backup id
            -n <area_id>    Delete files from <area_id> which have no numbers
                            in filenames.
            -o              Delete all offline files without backup id
            -p              'Purify' current dir
            -s              Simulate only, don't delete anything
            -t              Remove trash (files with file_area=-1)
            -z              Zuper slaughter, immense
            -v              Verbose execution
    
    -a : Deletes those files from <area_id> which exist on
         other areas. Works on filename basis only.
        
    -d : Deletes ALL files from the current dir which are found in the
         database. Don't use this on actual pimppa filearea dir.
        
    -f : Physically deletes all files which have failed the integrity
         check. Files will be marked as FLAG_OFFLINE and replaced later
         if the file is re-encountered e.g. by "pleech".

    -i : To be used after a successful backup operation. It deletes all
         files associated with a given backup_id, and gives them an
         FLAG_OFFLINE status to the database.

    -n : Delete files from <area_id> which have no numbers in filenames

    -p : Checks the current directory (which shouldn't be a filearea!)
         against your current file validity rules and assign patterns.
         Files which have negative assign patterns or violate the
         rules (e.g. NEEDNUMBER, see "src/pimppa.h") will be deleted.

    -s : Simulate execution only, don't change anything. Most useful
         with '-v'.
        
    -t : Deletes file-entries from database having file_area=-1. Such files
         can occur if you move files to *DOOM* instead of deleting them.

    -z : Deletes offline files from areas marked AREA_INCOMING,
         creates negative assign patterns for ALL of those files,
         and deletes the files from the database.

    pdesc
    -----
    Used to give text descriptions to files.
    
    Usage: pdesc <files> <"Desc">

           <files>          The files you wish to describe in current dir
           <"Desc">         The description you wish to give them

    e.g. 
    shell> pdesc goat*.jpg "Happy Goats"


    pf
    --
    "Pimppa Find". Searches for filenames from the database. 
    Requires SQL style wildcards (use '% for *' and '_ for ?'). 

    Usage:  pf [-p] <pattern>

            <pattern>       Filename pattern
            -p              Print pathnames as well

    The printing format is '[path/][filename] [backup_id if nonzero]'.
   
   
    prm
    -----
    Deletes files from the filearea matching current dir. Both
    the physical files and the records in the p_files -table 
    will be deleted. A new assign pattern will be created 
    for each file having -1 as the destination area, so that 
    all similar files will be killed when encountered in 
    the future.

    Usage:  prm [-v] <files>
  
            -n              Nice kill. Don't create negative assign
                            patterns or delete the file entry from database, 
                            just mark the file as offline and delete the
                            actual file.
            -v              Verbose execution

    Use prm with caution. :---)

    Note: If you wish to keep the MD5 checksums of the files,
    preventing the renames of the same files entering your
    system again, you should use "prm -n" or move the files 
    to area *DOOM*, e.g. "pmv <files> *DOOM*".


    pleech
    ------
    This is the command line downloading tool. It fetches files from 
    the newsgroups or ftp's you have defined. In case of news,
    promising articles are downloaded, then decoded, duplicates are
    re-checked, and suitable files are assigned to their appropriate 
    fileareas if they match an assign pattern. Otherwise they will 
    be entered to the default destination area. FTP-downloading
    is similar.

    As default, pleech uses the NNTP server found in p_misc.

    Usage:  pleech [options]
    
            -a <n>          Wait After n messages, def=10
            -f <urlfile>    Read (ftp-url, target area name) pairs from file,
                            leech the url recursively
            -g <pat>        Newsgroup name SQL pattern to match, def=all
            -H <host>       NNTP server name.
            -l              Lenient article selection while leeching news
                            (leech everything, select files after decoding)
            -n              Nazi behaviour. Deletes all incoming files which 
                            don't match an already defined assign pattern.
            -U <username>   NNTP server username, if needed.
            -P <password>   NNTP server password, if needed.
            -q              Quiet operation (don't print newsleech BPS etc)
            -v              Verbose output
            -w <secs>       Wait secs after [-a n] messages, def=2

    pleech depends on "suck" and "uudeview" for news and "fget" 
    for ftp. They must be on command execution path.
    NOTE: fget *NEEDS* to be patched (patch in "patches/" dir)!
    
    Newsleech example:

    shell> pleech -v -g %supreme.moose%
    Verbosely downloads all active newsgroups containing 
    string "supreme.moose" in the group name.
   
    FTP use:
    
    With -f, the urlfile is a text file containing one "URL AREA_NAME"
    pair per line, e.g.

    ftp://ftp.gigaturd.org/pub/pics incoming
    ftp://user:pass@grindcore.com/core incoming2


    pmarkoff
    --------
    Goes through your fileareas and marks offline files
    (not available in their directories). Such files are 
    skipped by most operations, like viewing and backuping.

  
    pmd5sum
    -------
    This is used to create/print/lookup/scan MD5 checksums of
    the files in system. Some of the operations below affect 
    only online files.

    Usage: pmd5sum <options>

    -b <id>   Create md5sums for backup id <id>
    -c        Create md5sums for online files
    -l <sum>  Look up files matching a given hex MD5 sum
    -p <pat>  Print sums of files matching an SQL filename pattern
    -s        Scan whole database for collisions
    -v        Verbose operation
    -w        Wipe DB from MD5-colliding files (in the case of
              collision, the oldest file is kept).

    Other usage is self-explanatory except that of -b. If 
    you use it, you must either move the files back online,
    or the easiest way is that if you have backuped the
    stuff using pbackup, and all your fileareas reside under
    "/stuff", do something like "mv stuff old ; ln -s /cdrom /stuff",
    then do "pmd5sum -b <id>" and after that replace the
    softlink "/stuff" with the original directory.

    
    pmv
    ---
    PIMPPA mv. Moves files from the area matching current dir 
    to a destination area. The matching area will be found out
    by using getcwd() and p_areas -table.

    Usage:  pmv <files> <destination_area>

    Example:
    shell> pmv fish*.jpg aquarium
    
    would move fish*.jpg to a filearea called 'aquarium'. The actual 
    destination directory location is fetched from the database.


    pnewarea
    --------
    Adds a new filearea to the database. Note that you must create
    the actual directory yourself with "mkdir". 

    Usage:  pnewarea -n <name> -p <dirpath> [options]

            -a             Mark area as AREA_NOASSIGN
            -i             Mark area as AREA_INCOMING
            -I <id>        Suggest id number for the new area
            -n <name>      Area name, MUST
            -t             Mark area as AREA_NOTRANS
            -T             RegExp dupecheck pattern, default $
            -p <dirpath>   Area directory path, MUST

    Example:
    shell> pnewarea -n pasture -p /stuff/pasture/
    
    would add a new filearea with name "pasture" and 
    directory path as "/stuff/pasture/". 


    pnewgrp
    -------
    Adds a new newsgroup to the system. Every group needs some 
    filearea as a default destination for the decoded files.

    Usage:  pnewgrp -d <dest_area> -n <group_name>

            -d <dest_area>  Destination area name, MUST
            -n <group_name> Newsgroup name, MUST


    ptest
    -----
    File integrity checker. Searches for untested files and
    tests them. Sets the database flag "file_integ" accordingly.
    The utils used for integrity checking of different filetypes 
    are defined in PIMPPA SQL table "p_types" and can be 
    modified from the GUI "bowser". If column "type_testokstr"
    is an empty string, the test command return value will
    be used instead. 0 == FILE OK, anything else == FAILED.

    Usage:  ptest [options]

            -n              Nazi behaviour. Delete files which fail
                            the integrity check. Use with CAUTION.
            -v              Verbose execution

    Note: "sql/example_types.sql" -file has example configs 
    for some integrity checkers and file transformers.
    

    ptrans
    ------
    File transformer. This can perform tasks like file optimization,
    conversion or perhaps add/remove spam (yuk). It searches for
    integrity-test passed files which have not yet been transformed 
    and tries to transform them. The utils for transformation are 
    defined in PIMPPA SQL table "p_types" and can be modified from 
    the GUI "bowser". 
    
    If some filearea has "area_flags" set to AREA_NOTRANS
    or AREA_INCOMING, the contents of the area will be skipped.

    Usage:  ptrans [-v]

            -v              Verbose execution

    Notes: 
    1) Successful transformation will usually modify the
       md5-checksum of the file.
    2) "sql/example_types.sql" -file has example configs 
       for some integrity checkers and file transformers.


8. Scripts
----------
These are just shell scripts, so you can easily edit 
them with a text editor to suit your needs.

Note: the viewing scripts are currently for picture material,
and default to "gqview" as the picture viewer. gqview is quite nice. :)
All the viewing scripts affect all files anywhere on the 
PIMPPA system (unless they are offline or on incoming areas). 


    p_areas
    -------
    Lists all your fileareas, their paths and their area ID numbers.
    
 
    p_con
    -----
    Prints out the contents for a given backup volume id
    in "Area | Megs" format.


    p_groups
    --------
    Prints all newsgroups and their target areas.
    
    Usage: p_groups <options>

            -a      Print only active groups
            -d      Print only disabled groups

    p_gtog
    ------
    Toggle active/disabled status of newsgroups matching
    given SQL newsgroup name pattern. Disabled newsgroups
    are not leeched by pleech or bowser.

    Examples: 
    shell> p_gtog %humbug% (toggles all groups with humbug in the name)
    shell> p_gtog alt.test (toggles just group "alt.test")
    

    p_loc
    -----
    Finds out backup id's containing files from a given filearea,
    specified by area name (or sql wildcard containing pattern).


    p_maint
    -------
    Useful to run daily from crond after "pleech". It just 
    performs "padopt", "ptest" and "ptrans". 


    pv_desc
    -------
    Views files matching a given SQL format file description. 


    pv_last
    -------
    Views files that arrived since the last run of this script
    (Bowser's "Extras/View since last" just executes this script.)


    pv_name
    -------
    Views files matching a given SQL format filename pattern.

    E.g. to display your kitten collection:
    shell> pv_name pussy%


    pv_since
    --------
    Views files which are newer than given number of days.


    pv_sql
    ------
    Views files by any suitable SQL WHERE statement. MySQL 
    manual is a suitable starting point if you don't know SQL.


    rc2sql
    ------
    Converts and inserts a suck .rc file (newsgroup list)
    to 'p_groups' table.


    viewdeep
    --------
    Actually not much to do with PIMPPA system. If you have used
    'wget' to mirror some website, but do not bother to click around
    the zillion directories, you can for example use 'viewdeep "*.jpg"'
    to give you a quick access to all jpg files in the current dir 
    and all its subdirs.


9. Some behaviour notes
-----------------------
How it works? What it eats? 


9.1 Assign patterns
-------------------
"pleech" (and bowser Leech, which uses the same routines)
decides based on assign patterns where it should store 
the files whose filename matches some pattern in the 
database. 

For each filename in the system there can be a (pattern, dest_area)
pair in "p_assign" -table, telling where similarly named files should 
be moved in the future. Naturally all files belonging together 
should map to the same pattern for this scheme to do any good.

The pattern is constructed from the filename as follows:

1) All numbers are converted to character '0'.
2) Letters after the last number and before the last dot ('.')
   are considered as indexes, and converted to '1' IF there's
   no more than two of them.
3) After last '.', all alphabetical letters stay intact.

E.g. filename       =>  pattern
     --------           -------
     "ab-103-h.jpg" => "ab-000-1.jpg"
     "ab-115-z.jpg" => "ab-000-1.jpg"
     "ab-ccc-1.jpg" => "ab-ccc-0.jpg"
     "ab-1-1ab.jpg" => "ab-0-011.jpg"
     "ab-1-def.jpg" => "ab-0-def.jpg"
     "ab-01a-2.jpg" => "ab-00a-0.jpg"

The assign patterns are by no means foolproof. One reason
is different files being created around the world with same 
names. However, its fairly good with really big series 
having some uncommon filename prefix like "gwo-bah-???.zip".
But it fails with files named imaginatively like 
"image001.jpg" which surface on every corner.
    
Example: If you have a pattern "bozo_000.png" pointing to area 5, 
"pleech" would send files named "bozo_123.png" and "bozo_124.png" 
to area 5, but files "bozo_abc123.png" and "bozo_100.jpg" would 
end up on the default destination area.

PIMPPA utils like "pmv", "padopt" and Bowser automatically 
update and create assign patterns, and the whole pattern 
database can be reconstructed with "passign".

Assign patterns are not created or used for areas marked
as AREA_NOASSIGN or AREA_INCOMING.

NOTE: Special destination area (a_dest) values:

-1  :  A negative assign pattern destination area id will cause
       all matching files to be quietly discarded by "pleech" and 
       Bowser in the future. In english: KILL the matching files. 
 0  :  Destination 0 means that the particular pattern is 
       disabled and won't be used. The pattern won't be replaced i
       by pimppa when matching files are moved or adopted. 
       Sorting those files will be left to the user.
>0  :  Some normal filearea.

The value 0 must be set by hand, e.g. in case you notice some
particular pattern causing incorrect classifications all the time.


9.2 Miscellaneous
-----------------
By default, "pleech" and bowser leech convert 
all filenames to lowercase, and discard all

1) duplicate files (see 9.3 RegExp dupecheck) 
2) MD5 -checksum colliding files (see 9.4)

For better spam avoidance, you should probably configure 
pimppa to discard

3) files that have no numbers in the filenames        (CFG_NEEDNUMBERS)
4) files that have only numbers before the extension. (CFG_NEEDOTHERS)
5) files that have whitespace in the filenames        (CFG_NOSPACE)

Usually such files are renames, spam or just plain nuisance. 

Use "pcfg" or "bowser" to change the settings to your
liking. See also 5.6, Changing Preferences.

Some additional behaviour options may be added in the future,
if some useful come to mind. I'll happily receive all 
suggestions and ideas!


9.3. RegExp dupecheck
---------------------
As a default, pimppa leech checks based on filenames that 
incoming files do not already exist on any filearea. 
If they do, the incoming counterparts are called duplicates 
and deleted (unless the already existing files have failed 
the integrity check - in that case they are replaced with
the new ones).

You may wish to relax the duplicate checking to check 
only from particular areas.

Example case.

You get "goat100.jpg" from "alt.binaries.pictures.animals",
and later a file with the same name from "alt.worship.goatlord". 
Now there's a good change these are not the same files. You
might prepare for cases like this by relaxing the dupechecking
as follows:

Group: "alt.binaries.pictures.animals"
        => default destination area "0animaltmp"
            => set "0animaltmp" to dupecheck from areas "cats", "dogs", "sheep"
Group: "alt.worship.goatlord"
        => default destination area "0occulttmp"
            => set "0occulttmp" to dupecheck from area "weirdstuff" only

The dupechecking is set for the destination areas, not for the 
sources themselves. (E.g. many newsgroups may map to the same 
destination area and follow the same dupecheck patterns).

To set RegExp duplicate checking, just set a proper area_id
RegExp pattern for any incoming area (modify 'area_targets' -column). 

Examples:

$               Default, check from all areas
^1$             Dupecheck only from area with ID 1
^7$|^10$|^15$   Dupecheck from areas 7,10,15.

You can set these from "bowser" or directly by "mysql".

NOTE: RegExp duplicate checking also affects assigning by
leech operations: only those assign target areas are seen valid which
are matched by RegExp. If there is no match, default destination 
area is used. If assign target is negative (*DOOM*), the file 
will be deleted (is this behaviour is wise?), no matter what 
RegExp says.


9.4 MD5 -based dupecheck
------------------------
For each file entered to pimppa system, a 128bit MD5sum
will be calculated. The sum is compatible with RFC 1321. 

If STRICT_MD5 is used (as default), pimppa utilities
will delete all incoming files which have an md5sum
colliding with some existing md5sum. This gives a really
good duplicate discarding system, though some innocent
files might be deleted because of false checksum collisions. 
After going through my database I didn't find such a case,
but over hundred "valid" collisions (which were renames: 
exactly same file, but with a different filename).


10. MySQL table explanation
---------------------------
Main database is "pimppa" and it's owned as by user "pimppa".

"p_areas" is the table containing all your fileareas.

    area_id         Unique area ID number
    area_name       Area name. Should be logical and quick to type.
    area_path       The directory path for the files of this area.
    area_flags      Properties of this area (hints for utils)
    area_targets    RegExp pattern for dupechecking incoming (see 9.3)

"p_assign" contains the assign patterns - where "pleech", "bowser"
and "padopt -m" should deposit certain files. Negative destination
area makes utils to delete the incoming file. Zero destination means
that this pattern is disabled.

    a_pattern   The filename pattern to match
    a_dest      Destination area_id for all matching files

"p_files" contains all the files you have.

    file_id     Unique file ID number
    file_name   File name, unique per filearea.
    file_size   File size in bytes
    file_area   Area ID where this file should be
    file_integ  File integrity check status
    file_trans  File transformation status
    file_date   The date when you got this file
    file_backup The ID of the backup, 0 if none.
    file_desc   Optional ASCII text description of this file
    file_flags  Is there something special with this file (offline?)
    file_md5sum 128 bit MD5 checksum for this file
    
"p_groups" contains information about newsgroups.

    g_name      Unique name of the newsgroup, e.g. "alt.binaries.test"
    g_last      Last msg read -pointer
    g_flags     Newsgroup flags (can be |= GROUP_DISABLED)
    g_dest      Destination filearea id
    
"p_misc" is a really general table for various PIMPPA-utils
to store their status information and configuration.

    misc_key    Identifying unique key for the info
    misc_data   The actual data associated with the key

"p_types" contains the information how to handle various filetypes.
If type_testokstr is an empty string, the test command return 
value will be used. 0 == FILE OK, anything else == FAILED.

    type_ext        File extension of this type, e.g. "JPG".
    type_testcmd    Command to use to test a file of this type 
    type_transcmd   Command to use to transform a file of this type
    type_testpos    Position of the success string for testcmd (deprecated)
    type_testokstr  The actual "all correct" string given by testcmd.
    type_transto    Destination file type when transforming.


11. Feedback
------------
Email <iwronsky(-at-)users.sourceforge.net>. Any bug-reports, 
patches, comments or suggestions are welcome. Especially if you 
have an idea about some new functionality/tool/script that would 
benefit PIMPPA, don't hesitate. And if you know how to actually do it,
all the better!

If you prefer to use PGP, my public key is at the end of 
this README.


12. Glossary
------------

    "*DOOM*"
    --------
    This is a point-of-no-return -filearea which should exist
    on all systems. Its area_id is "-1" and area_path "/dev/null".
    All files moved or assigned to *DOOM* will never be seen again.

    MD5 -checksums of the moved files will be kept but the 
    files assigned there in the future won't leave a trace.


    "assign pattern"
    ----------------
    Downloading decides based on assign patterns where files
    should be stored whose filename match an assign pattern in the 
    database. Section 9.1 tells how the patterns currently operate.
   

    "filearea"
    ----------
    PIMPPA is structured so that every file is on a certain 
    filearea. Fileareas are created with "pnewarea". All files of 
    similar content should be on a certain filearea, so you can 
    easily find them. It's quite like a normal directory, 
    except that some additional info of the filearea contents
    is kept in the p_files database to make the backup, duplicate-
    check, lookup, etc, operations possible and fast.


    "Incoming" (AREA_INCOMING)
    --------------------------
    Fileareas to which "raw" material from newsgroups is decoded
    to. Incoming area contents will not be backuped or transformed. 
    
    In ideal operation pimppa is like a sorting network,
    the files can be seen travelling like this:

    Newsgroup  Filter 1        Filearea     Filter 2         Filearea
    
    group1-| 
           |->-[autofilter]->- Incoming1 ->-[humanfilter]->- Quality1
    group2-|      |  |                         |                |
                  |  `------------>--------------->-------------'
                  |                            |
                  Discard                      Discard

    etc. Due to the assign patterns, recognized files can be
    moved to correct fileareas without human interaction. Also
    the duplicate checking (filename and md5check) and certain
    requirements for filenames allow some incoming files to 
    be discarded automatically.
    

    "transform"
    -----------
    Operation which can be performed (once!) for a file of
    a certain filetype. This operation can be a conversion
    of ".GIF" to ".PNG", an optimizing of a JPEG, or whatever.
    You can specify the transformation command and result
    filetype (which can be same as source type) in 'p_types'-
    table.

    Some example transformations are in "sql/example_types.sql"


    "offline"
    ---------
    Files which exist on some filearea but not in its respective
    directory are called offline. The files may have been backupped 
    and deleted, or just lost. Most PIMPPA utilities skip offline 
    files. Files which are present are called "online". 


    "PIMPPA"
    --------
    PIMPPA is a fabulous content seeking/devouring creature 
    or monster in the forgotten scandinavian mythologies.
    

13. PGP
-------
This is my PGP public key. Encrypted mail would be nice,
not because there was something to hide, but because privacy
should be considered normal and available for anyone,
not a limited delicacy for the few and the weird. Of course
giving a key here doesn't guarantee much - it could
belong to anyone. ;)


-----BEGIN PGP PUBLIC KEY BLOCK-----
Version: GnuPG v1.0.1 (GNU/Linux)
Comment: For info see http://www.gnupg.org

mQGiBDawXNERBADWkpWZoDuB//oUsvlmCKNufobFKdTX1SGfqDffoHgOO+01LVU9
QZl+snoSoAz7TkRnEep34mdRcx8IRe/xuLi88jEiI2zVLCHoYFBDB6lpxfSGYjJq
IfqCxv1GKMTigykLI5oxyLgYyrONL00uZzPRuTnSrwZqWuqqoHANlxed1QCg/3Xy
FLgvmPRZkBVZ7cjAV+ZlNEEEAKNFmfPGTwKd5X0MBSI4aIUXKO3G+Narjn9vVDoi
LsU41LR3VO50UeWvnbSEi93YTDwXpFuKK4nMKG1Z5T8pZPQ6NqcOu6NxBm6GHMoR
L6ZI+oRgbdAG2shTlfbnZ10qyPhTg0VDf+zfD7sx/vK7jo5uywrmiTUvQvEli4ps
6KCFBADGSC6Uf/5lNB7+4VAko0j1G6m3cnOekvl6/XX+FihPob9NECZbiJR3lvzR
ayuJaOSoX/zNJEHDzfOu1qzjOvbWBWY+JBkJ1QYvpB1A+Y84QFcBJRlrvTcXxgfv
urMkptlO3yUY7UPKIt7+52XTAesfWZkmkgC9cPQ5gD2l1FjyZbQhSWdvciBXcm9u
c2t5IDxpd3JvbnNreUB5YWhvby5jb20+iEsEEBECAAsFAjd9HuMECwMBAgAKCRAg
/QqD5XhjksYfAKD8QRg9EPsiLacwNPTGppusFCCvpQCg6k0tB7FVJaLjunbWKItZ
LuPHvFm5Ag0ENrBc0hAIAPZCV7cIfwgXcqK61qlC8wXo+VMROU+28W65Szgg2gGn
VqMU6Y9AVfPQB8bLQ6mUrfdMZIZJ+AyDvWXpF9Sh01D49Vlf3HZSTz09jdvOmeFX
klnN/biudE/F/Ha8g8VHMGHOfMlm/xX5u/2RXscBqtNbno2gpXI61Brwv0YAWCvl
9Ij9WE5J280gtJ3kkQc2azNsOA1FHQ98iLMcfFstjvbzySPAQ/ClWxiNjrtVjLhd
ONM0/XwXV0OjHRhs3jMhLLUq/zzhsSlAGBGNfISnCnLWhsQDGcgHKXrKlQzZlp+r
0ApQmwJG0wg9ZqRdQZ+cfL2JSyIZJrqrol7DVekyCzsAAgIH/AzQXkgMRpwsbEVh
XSEH/5kbN4Ls9LbFMkPelmaODl2W2wjmWa+7loBFnKn+9WHh77/GLMzHGPYoTzZv
wp6bAYbcq4cu20qdW2tTIfUXJz+ey3r5rwFR5y5qkiBqfFczepY0biUcUI7dWt/Q
LUyN6oVyVAjclmfvA/JWi7LmMRl6Jo1doKXLYhHOuFXkoGqIExrO9EUKTMGsa0Lm
uJVv6kb0v9EAyiJU/zzMvKotPtdzzPqz2m+0mt/XsMhfbT6xl2XkmESvQhgev6Yh
DYpzVSZOZeZ7Etzpp2eDwfP4AU23ge6KFO7g33cSEJilBe7x3ZkiTb5Hgqs3FnWi
t9EqwDmIPwMFGDawXNIg/QqD5XhjkhECewUAn3P1gtt1Y3DZRRWvJ9TgNCtc+qcp
AJ9p6oceyAzcCw87KNm3kW7u6gBK6g==
=Hq/M
-----END PGP PUBLIC KEY BLOCK-----


<EOF>
