Table of Contents
Overview
Why? (Long Term Ivory Tower Vision)
Why? (Short Term Practical Vision)
Using Aware
Features
Anatomy of an Aware Daemon
Overview
The Aware project is an
effort to create a software framework to measure, monitor, and control computer system resources.
Aware is intended to enable system administrators tune
system variables, set monitoring/security alarms and build adaptive distributed systems.
Aware modules may be linked into applications making them 'aware' and able to
participate in the larger managed system.
Ultimately, multiple systems will monitor themselves and others, cooperating
to make decisions to optimally tune performance, proactively enhance security and
compensate for faults.
The Aware software is a high performance distributed event processing framework built for
systems management. It comes with probes for the common network services
and system resources. Additionally, Aware allows the cross-correllation of many different
streams of information.
The development goals are:
- High performance with small footprint to minimize the impact on host and allow deployment on small systems.
- Dynamically extensible such that 3rd parties can supply modules outside the Aware core software base.
- Linkable libraries such that applications may incorporate Aware features and cooperate in system control and monitoring.
- Portable, clean and simple code.
The development has 3 major phases:
- Resource measurement and monitoring tools
- High speed efficient event processor
- Linkable library APIs
- Rich set of monitoring probes and notification/action handlers
- High level probe/alarm/handler configuration language
- Standalone applications for system monitoring/analysis
- Network wide adaptive control of resources
- Secure messaging between Aware enabled applications
- Awareness
- Smart/Adaptive distributed cooperative system control
Initially, the development will be confined to Linux.
However, it is planned that the system will be quickly be ported other flavors of Unix
(e.g., Solaris, the BSDs, OSX/Darwin, etc.)
Aware is implemented in ANSI C and is written to be as portable as possible.
Why? (Long Term Ivory Tower Vision)
It is not uncommon today to have a system with more than 100 or even a 1000 networked computers.
Since the complexity of managing such systems rises exponentially with the number of computers
connected (i.e., consider multiple ways each computer can interact with itself and others) we
will rapidly reach a point in the next 5 years where it will not be possible to manage
our vast networks of computers using the human intensive methods we generally employ today.
The solution to this problem lies within the systems themselves. In order to manage the
next generation of computer systems that connect millions of computers, the computers
themselves must take on the burden of managing such issues as security and availability.
That's not to say the goal is to eliminate the human element in the decision making processes,
rather, the systems need to be self monitoring and adaptive such that they only present
the humans a manageable amount of information. IBM has a high level overview of the elements
required to create such a system, they call these the 8 Elements.
The goal of this project is to build a working framework on which such systems
can be built.
Why? (Short Term Practical Vision)
In the real world, conceptual systems are of little value. A goal of the Aware project
is to make the software immediately useful while building on a framework that
can extend to reach the loftly goals outlined in the previous section.
In order to build smart/adaptive feedback controlled distributed systems 4 major
areas of functionality are required:
- Sensors:A comprehensive set of sensors that gather relevant information
- Analyzers: Components that process data from the sensors and issue controller commands
- Controllers: Components that change system state (e.g., run programs, change system parameters, control devices)
- Stores: Components responsible for the persistent storage and retrieval of sensor data
Essentially, the above describes a system/network monitoring system (S/NMS). Therefore, the short term goal
of the Aware project is to build and extensible system/network monitoring platform. However,
this is being done in the context of the long term vision and implies an implementation framework very
different from most current monitoring packages. This framework has the following characteristics:
- Open source implementation allows for robust code base and customization
- Common core engine implements a model of event processing objects called probes and handlers
- A "plug in" style mechanism allows dynamic addition of probes and handlers
- Agents are composed of a set of running probes and handlers
- Agents can get their configuration from other agents (e.g., a centrally managed set of agent configurations)
- Agents can communicate with other agents using connection oriented, connectionless and broadcast based methods
- Agents have authentication/authorization mechanisms
- High performance and low impact implementation allows agents to do their jobs very efficiently
Once Aware has a comprehensive set of system management functions it can be extened in a natural way
to the less mundane issues such as proactive security, adaptive provisioning, etc.
Using Aware
Software Architecture
The Aware software architecture has 3 conceptual levels, each with a
different target user base:
- API: C code.
Intended User Base: Programmers needing a high
performance async event engine
- API: Wire scripts and commandline tools.
Intended User
Base: Sysadmins that don't mind scripting
- API: Web based UI w/database.
Intended User Base: Sysadmins
that don't want to script and/or want easy to use tool.
Level 1 allows a programmer to create programs that use the aware event
engine for high performance async event
oriented applications. The 'discover' program that comes with the
distribution is a port scanner written in C that
uses the event engine. Programmers can write their own probes and
handlers in C and they will be available
in C and (if they follow the standard API) via wire scripts. Programmers
can develop these probes and handlers
outside the aware distribution because each probe/handler is a dynamic
object, loaded at run time.
Level 2 makes it very easy to string together probes and handlers in
'wire' scripts. Using the 'exec' probe/handler you can
integrate with commandline tools (e.g., usb-snmp tools). The examples
directory in the
distribution contains many working examples of 'wire' scripts.
Level 3 builds on the previous levels by being a web based application that uses a
master 'wire' script it generates from data in the database. This
application makes basic networking
monitoring very easy to set up. Also, you can integrate custom wire
scripts local or remote from the machine
running this application.
Aware Agents for Network Monitoring
Agents may be trivially configured to monitor all standard IP based services. Typically,
you would install an Aware agent on a small number of machines inside your network that
monitor all servers, sending notifications when a service or server is unavailable. Agents
running outside your network monitor your system's availability from other
networks or the Internet.
Aware's support for UPnP allows monitoring all UPnP computers (e.g., Windows XP w/Internet Connection Sharing) and devices.
Aware Agents for System Monitoring
Aware can monitor all standard system statistics (e.g., processes, load, swap, free memory) generating
events when system administrators require notificaction or taking actions (e.g., restarting
failed applications).
Aware Agents for Security Monitoring
Aware can monitor all login/logout by users and watch sensitive files for modification.
Aware can work with any security application that generates text logs, generating alerts
when there are indications of a security issue (e.g., watching the tail of /var/log/messages
for security related messages like 'su root' or dropped packets by a packet filter).
Aware's high speed event processing make it ideal to work with IDS's like
.
Aware Agents Cooperating
In the future, Aware agents will be able to communicate securely allowing feedback based control
of your system. For example, Aware agents may detect a drop in traffic on your
web servers (e.g., at night EST) and send events to a central Aware agent that
calculates the aggregate load and the optimal number of servers. This agent then
sends shutdown messages to the uneeded servers (we assume that your load balancer
will gracefully reroute traffic to the remaing servers), thus saving considerable money
on power costs. Later, as traffic increases, the central Aware agent sends a message
to a power sentry on servers that need to be brought online to service the traffic.
You can implement a dynamic security policy using Aware agents. Consider a system
with Aware agents monitoring sensitive files and network traffic on many servers and
that one of these agents detects an active network based security breach. This agent
may send a notification to the Aware agent on the firewall to immediately add a firewall rule to notch out all
traffic from the offending IP address.
Features
Easy to use and install
Aware based programs read a text file that describes the probe and handler configuration.
A menu based configuration tool enables rapid deployment.
For example, here is how you might set up a probe to monitor your
web server, testing for connection timeouts, 404 errors, and application Oracle database errors:
// define events
set noconnect create event { name: "noconnect" priority: 1 }
set four-oh-four create event { name: "file not found" priority: 1 }
set oracle-error create event { name: "oracle error" priority: 1 }
// bind a list to save on typing
list sa-events $noconnect $four-oh-four
list dba-events $oracle-error
list allevents $sa-events $dba-events
// create the probe, check every minute
create probe http {
url: http://webserver.mydomain.net
noconnect: $noconnect
match: $four-oh-four "" "HTTP/.* 404" "webserver home page is missing!"
match: $oracle-error "" "ORA-[0-9][0-9][0-9][0-9][0-9][0-9]" ""
timeout: 5
cycletime: 60
}
|
Here is the definition of 4 handlers that register to receive events from the above web server
probe. The first registers to receive the noconnect events and sends a page
to the system administrator. The second registers to receive the noconnect and 404 events and sends email
to the system administrator. The third registers to receive the application data base
errors and sends email to the DBA on call. The forth handler registers to receive
all the events and logs them to a file locally as a permanent record.
// Send page alerts to sysadmin for connection failures
create handler execp {
cmd: "kermit /usr/share/mon/pagesysadmins"
regevent: $noconnect
}
// Send email alerts to sysadmin for connection failures and 404s, these
// usually indicate a network or configuration problem
create handler execp {
cmd: "mail -s \"alert: \$n\" sysadmin@mydomain.net"
input: "Aware alert:\n \$n \$v\n\n"
regevent: $sa-events
}
// Send email alerts to dbas for Oracle errors
create handler execp {
cmd: "mail -s \"alert: \$n\" dba@mydomain.net"
input: "Aware alert:\n \$n \$v\n\n"
regevent: $dba-events
}
// Record locally. Rotate log files daily
set logger create hlogger { filename: /log/webalerts rotate: daily }
create handler log { logger: $logger regevent: $allevents }
|
The distribution comes with many working examples that will allow you to rapidly
configure your system.
Probes
Probes run on a predefined schedule measuring system statistics and probing network services.
Probes generate events that handlers may act on (e.g., if the httpprobe detects Oracle errors
within a Web based application it will generate an event that triggers a handler to send
a notification and another handler to shutdown the web server).
Currently implemented probes:
- Network Service Availability
- System Monitoring
- Communications
- Database Integration
- Other
Handlers
Handlers receive events from probes or other handlers and take actions
such as sending notifications, managing processes and computing statisitics.
Currently implemented probes:
- Notification
- Database Integration
- Statistics
- Filtering Events
- Control
- Communications
- Other
Anatomy of an Aware Daemon
The core of the Aware infrastructure is an event processing monitor composed
of 3 major entities: probes, receivers and handlers. The
probes and receivers generate events that are processed by
the handlers (which may, in turn, generate events for other
handlers).
The system comes with a rich set of probes, receivers and
handlers. Users may extend the system with their own. A declarative
language may be used to configure a daemon. The software comes
with many examples.
The probes, receivers and handlers run within a central monitor. All
execution is asynchronous to allow large numbers of these entities to
run simultaneously without significant impact on system performance.
The central monitor implements a form of cooperative multitasking
where entities are expected to do their processing very quickly or
yield and resume processing later. This was done to allow the monitor
to scale to many thousands of probes, receivers and handlers. Entities
that require long running codes or use blocking system calls may
be configured to run within a "tasklet". A tasklet is a thread
within which a probe or handler may run. The use of tasklets
is discouraged, however, they are necessary if dependent external
libraries do blocking systems calls which would stall the execution
of the monitor (e.g., database client libraries).
Probes
An Aware probe is a compuational unit that runs periodically and may
generate events.
Examples:
- Timer Event Generator
- File Status
- Process Watcher
- Text Log File Watcher
- TCP Port Probe
- Ping Probe (ICMP)
- HTTP Probe
Receivers
Receivers are probes that have registered
to be signaled when data is received on a file descriptor(s).
Examples:
- Universal Plug-n-play (UPnP) events
- SNMP events
- Sensor Events (e.g., temperature)
- Aware Network Events
- Window System Events
Handlers
An Aware handler is a compuational unit that registers to receive
either specific events or classes of events based on a bit vector
pattern match
and may itself generate events.
Examples:
- Calculate Long Term System Stats
- Send Email Notification
- Send UPnP Messages
- Send SNMP Traps
- Send Aware Network Events
- Kill Processes
- Run Commands
- Control devices
|