Elegant Software



Table of Contents

Overview
Why? (Long Term Ivory Tower Vision)
Why? (Short Term Practical Vision)
Using Aware
Features
Anatomy of an Aware Daemon

Overview

The Aware project is an effort to create a software framework to measure, monitor, and control computer system resources. Aware is intended to enable system administrators tune system variables, set monitoring/security alarms and build adaptive distributed systems. Aware modules may be linked into applications making them 'aware' and able to participate in the larger managed system.

Ultimately, multiple systems will monitor themselves and others, cooperating to make decisions to optimally tune performance, proactively enhance security and compensate for faults.

The Aware software is a high performance distributed event processing framework built for systems management. It comes with probes for the common network services and system resources. Additionally, Aware allows the cross-correllation of many different streams of information.

The development goals are:

  • High performance with small footprint to minimize the impact on host and allow deployment on small systems.
  • Dynamically extensible such that 3rd parties can supply modules outside the Aware core software base.
  • Linkable libraries such that applications may incorporate Aware features and cooperate in system control and monitoring.
  • Portable, clean and simple code.

The development has 3 major phases:

  1. Resource measurement and monitoring tools
    • High speed efficient event processor
    • Linkable library APIs
    • Rich set of monitoring probes and notification/action handlers
    • High level probe/alarm/handler configuration language
    • Standalone applications for system monitoring/analysis
  2. Network wide adaptive control of resources
    • Secure messaging between Aware enabled applications
  3. Awareness
    • Smart/Adaptive distributed cooperative system control

Initially, the development will be confined to Linux. However, it is planned that the system will be quickly be ported other flavors of Unix (e.g., Solaris, the BSDs, OSX/Darwin, etc.) Aware is implemented in ANSI C and is written to be as portable as possible.

Why? (Long Term Ivory Tower Vision)

It is not uncommon today to have a system with more than 100 or even a 1000 networked computers. Since the complexity of managing such systems rises exponentially with the number of computers connected (i.e., consider multiple ways each computer can interact with itself and others) we will rapidly reach a point in the next 5 years where it will not be possible to manage our vast networks of computers using the human intensive methods we generally employ today.

The solution to this problem lies within the systems themselves. In order to manage the next generation of computer systems that connect millions of computers, the computers themselves must take on the burden of managing such issues as security and availability. That's not to say the goal is to eliminate the human element in the decision making processes, rather, the systems need to be self monitoring and adaptive such that they only present the humans a manageable amount of information. IBM has a high level overview of the elements required to create such a system, they call these the 8 Elements.

The goal of this project is to build a working framework on which such systems can be built.

Why? (Short Term Practical Vision)

In the real world, conceptual systems are of little value. A goal of the Aware project is to make the software immediately useful while building on a framework that can extend to reach the loftly goals outlined in the previous section.

In order to build smart/adaptive feedback controlled distributed systems 4 major areas of functionality are required:

  1. Sensors:A comprehensive set of sensors that gather relevant information
  2. Analyzers: Components that process data from the sensors and issue controller commands
  3. Controllers: Components that change system state (e.g., run programs, change system parameters, control devices)
  4. Stores: Components responsible for the persistent storage and retrieval of sensor data

Essentially, the above describes a system/network monitoring system (S/NMS). Therefore, the short term goal of the Aware project is to build and extensible system/network monitoring platform. However, this is being done in the context of the long term vision and implies an implementation framework very different from most current monitoring packages. This framework has the following characteristics:

  • Open source implementation allows for robust code base and customization
  • Common core engine implements a model of event processing objects called probes and handlers
  • A "plug in" style mechanism allows dynamic addition of probes and handlers
  • Agents are composed of a set of running probes and handlers
  • Agents can get their configuration from other agents (e.g., a centrally managed set of agent configurations)
  • Agents can communicate with other agents using connection oriented, connectionless and broadcast based methods
  • Agents have authentication/authorization mechanisms
  • High performance and low impact implementation allows agents to do their jobs very efficiently

Once Aware has a comprehensive set of system management functions it can be extened in a natural way to the less mundane issues such as proactive security, adaptive provisioning, etc.

Using Aware

Software Architecture

The Aware software architecture has 3 conceptual levels, each with a different target user base:
  1. API: C code.
    Intended User Base: Programmers needing a high performance async event engine
  2. API: Wire scripts and commandline tools.
    Intended User Base: Sysadmins that don't mind scripting
  3. API: Web based UI w/database.
    Intended User Base: Sysadmins that don't want to script and/or want easy to use tool.

Level 1 allows a programmer to create programs that use the aware event engine for high performance async event oriented applications. The 'discover' program that comes with the distribution is a port scanner written in C that uses the event engine. Programmers can write their own probes and handlers in C and they will be available in C and (if they follow the standard API) via wire scripts. Programmers can develop these probes and handlers outside the aware distribution because each probe/handler is a dynamic object, loaded at run time.

Level 2 makes it very easy to string together probes and handlers in 'wire' scripts. Using the 'exec' probe/handler you can integrate with commandline tools (e.g., usb-snmp tools). The examples directory in the distribution contains many working examples of 'wire' scripts.

Level 3 builds on the previous levels by being a web based application that uses a master 'wire' script it generates from data in the database. This application makes basic networking monitoring very easy to set up. Also, you can integrate custom wire scripts local or remote from the machine running this application.

arch

Aware Agents for Network Monitoring

Agents may be trivially configured to monitor all standard IP based services. Typically, you would install an Aware agent on a small number of machines inside your network that monitor all servers, sending notifications when a service or server is unavailable. Agents running outside your network monitor your system's availability from other networks or the Internet.

Aware's support for UPnP allows monitoring all UPnP computers (e.g., Windows XP w/Internet Connection Sharing) and devices.

Aware Agents for System Monitoring

Aware can monitor all standard system statistics (e.g., processes, load, swap, free memory) generating events when system administrators require notificaction or taking actions (e.g., restarting failed applications).

Aware Agents for Security Monitoring

Aware can monitor all login/logout by users and watch sensitive files for modification. Aware can work with any security application that generates text logs, generating alerts when there are indications of a security issue (e.g., watching the tail of /var/log/messages for security related messages like 'su root' or dropped packets by a packet filter).

Aware's high speed event processing make it ideal to work with IDS's like .

Aware Agents Cooperating

In the future, Aware agents will be able to communicate securely allowing feedback based control of your system. For example, Aware agents may detect a drop in traffic on your web servers (e.g., at night EST) and send events to a central Aware agent that calculates the aggregate load and the optimal number of servers. This agent then sends shutdown messages to the uneeded servers (we assume that your load balancer will gracefully reroute traffic to the remaing servers), thus saving considerable money on power costs. Later, as traffic increases, the central Aware agent sends a message to a power sentry on servers that need to be brought online to service the traffic.

You can implement a dynamic security policy using Aware agents. Consider a system with Aware agents monitoring sensitive files and network traffic on many servers and that one of these agents detects an active network based security breach. This agent may send a notification to the Aware agent on the firewall to immediately add a firewall rule to notch out all traffic from the offending IP address.

Features

Easy to use and install

Aware based programs read a text file that describes the probe and handler configuration. A menu based configuration tool enables rapid deployment.

For example, here is how you might set up a probe to monitor your web server, testing for connection timeouts, 404 errors, and application Oracle database errors:
// define events
set noconnect create event { name: "noconnect"  priority: 1 }
set four-oh-four create event { name: "file not found"  priority: 1 }
set oracle-error create event { name: "oracle error"  priority: 1 }

// bind a list to save on typing
list sa-events $noconnect $four-oh-four 
list dba-events $oracle-error
list allevents $sa-events $dba-events

// create the probe, check every minute
create probe http { 
 url: http://webserver.mydomain.net
 noconnect: $noconnect
 match: $four-oh-four "" "HTTP/.* 404" "webserver home page is missing!"
 match: $oracle-error "" "ORA-[0-9][0-9][0-9][0-9][0-9][0-9]" ""
 timeout: 5
 cycletime: 60
}

Here is the definition of 4 handlers that register to receive events from the above web server probe. The first registers to receive the noconnect events and sends a page to the system administrator. The second registers to receive the noconnect and 404 events and sends email to the system administrator. The third registers to receive the application data base errors and sends email to the DBA on call. The forth handler registers to receive all the events and logs them to a file locally as a permanent record.

// Send page alerts to sysadmin for connection failures 
create handler execp { 
 cmd: "kermit /usr/share/mon/pagesysadmins"
 regevent: $noconnect 
}

// Send email alerts to sysadmin for connection failures and 404s, these
// usually indicate a network or configuration problem
create handler execp { 
 cmd: "mail -s \"alert: \$n\" sysadmin@mydomain.net"
 input: "Aware alert:\n \$n \$v\n\n"
 regevent: $sa-events
}

// Send email alerts to dbas for Oracle errors 
create handler execp { 
 cmd: "mail -s \"alert: \$n\" dba@mydomain.net"
 input: "Aware alert:\n \$n \$v\n\n"
 regevent: $dba-events
}

// Record locally. Rotate log files daily
set logger create hlogger { filename: /log/webalerts rotate: daily }
create handler log { logger: $logger regevent: $allevents }


The distribution comes with many working examples that will allow you to rapidly configure your system.

Probes

Probes run on a predefined schedule measuring system statistics and probing network services. Probes generate events that handlers may act on (e.g., if the httpprobe detects Oracle errors within a Web based application it will generate an event that triggers a handler to send a notification and another handler to shutdown the web server).

Currently implemented probes:

Handlers

Handlers receive events from probes or other handlers and take actions such as sending notifications, managing processes and computing statisitics.

Currently implemented probes:

Anatomy of an Aware Daemon

The core of the Aware infrastructure is an event processing monitor composed of 3 major entities: probes, receivers and handlers. The probes and receivers generate events that are processed by the handlers (which may, in turn, generate events for other handlers).

The system comes with a rich set of probes, receivers and handlers. Users may extend the system with their own. A declarative language may be used to configure a daemon. The software comes with many examples.

The probes, receivers and handlers run within a central monitor. All execution is asynchronous to allow large numbers of these entities to run simultaneously without significant impact on system performance. The central monitor implements a form of cooperative multitasking where entities are expected to do their processing very quickly or yield and resume processing later. This was done to allow the monitor to scale to many thousands of probes, receivers and handlers. Entities that require long running codes or use blocking system calls may be configured to run within a "tasklet". A tasklet is a thread within which a probe or handler may run. The use of tasklets is discouraged, however, they are necessary if dependent external libraries do blocking systems calls which would stall the execution of the monitor (e.g., database client libraries).

Probes

An Aware probe is a compuational unit that runs periodically and may generate events.

Examples:

  • Timer Event Generator
  • File Status
  • Process Watcher
  • Text Log File Watcher
  • TCP Port Probe
  • Ping Probe (ICMP)
  • HTTP Probe

Receivers

Receivers are probes that have registered to be signaled when data is received on a file descriptor(s).

Examples:

  • Universal Plug-n-play (UPnP) events
  • SNMP events
  • Sensor Events (e.g., temperature)
  • Aware Network Events
  • Window System Events

Handlers

An Aware handler is a compuational unit that registers to receive either specific events or classes of events based on a bit vector pattern match and may itself generate events.

Examples:

  • Calculate Long Term System Stats
  • Send Email Notification
  • Send UPnP Messages
  • Send SNMP Traps
  • Send Aware Network Events
  • Kill Processes
  • Run Commands
  • Control devices



Aware | About | News | Examples | Download | Development Plan | Contact




(c) 2002 Russell Leighton, all rights reserved