README for DerMixD v1.6.0; date: 07.04.2008
also see web page at http://dermixd.de
*******************************

	Copyright (C) 2004-8  Thomas Orgis <thomas@orgis.org> and others
	See AUTHORS file for contributors.

	This program is free software; you can redistribute it and/or modify
	it under the terms of the GNU General Public License version 2 as
	published by the Free Software Foundation,

	This program is distributed in the hope that it will be useful,
	but WITHOUT ANY WARRANTY; without even the implied warranty of
	MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the
	GNU General Public License for more details.

	You should have received a copy of the GNU General Public License
	along with this program; if not, write to the Free Software
	Foundation, Inc., 59 Temple Place, Suite 330, Boston, MA  02111-1307  USA

For building and install notes see INSTALL. For the impatient: Run make, pick a target (such as gnu), run `make <target>` and wish yourself luck - and tell me if you don't have luck.

See CHANGES for main, user-concerning changes, ChangeLog for the developer view.

***************************

So, this is DerMixD, an audio mixing and network-listening daemon in the tradition of mixplayd (http://mixplayd.sourceforge.net) but done quite differently. It does basically what mixplayd does:

- play music files on a configurable number of input channels
- mix them together
- be controlled over a tcp (or UNIX domain socket)

But it does not do fading all of its own - you are supposed to define the fading volume and equalizer ramps via script commands (that's actually a _good_ thing!).
There's no explicit support for named pipes as output, but thinking about it, there is no special support necessary.
You do mkfifo, open that named pipe and let dermixd write raw audio to that filename. It doesn't need to _know_ that there is a named pipe involved (UNIX ftw!;-).

I wrote this mostly from scratch with some flexibility in mind using multithreading (what doesn't matter in terms of flexibility, but what matters nevertheless) and c++ with general input and output classes allowing to implement special classes for different audio formats or devices on both input and output side - hopefully - quite easily. 
Concernig the mp3 decoder input, the main difference is that dermixd acts as a mpg123 frontend to a current mpg123, keeping one decoder instance alive instead of restarting on every seek and load. I had already hacked this behaviour into mixplayd-0.55 but wasn't satisfied with the need to wait for several decoders one after another - that is the point where multithreading comes in: I wanted to wait for the I/O stuff in a parallel way. This was part of knowing then a file ended without watching mpg123's responses; I needed the timeouts and accurate info on track length... worked, but messy.
That said: The original mixplayd has no such problems since it doesn't act as a frontend. Mainly the frontend stuff brings complexity but it has befefits, too: No overhead through restarting mpg123 and reopening files all the time and interactive use of mpg123's equalizer.

Also the mixing itself is more flexible. Every input channel gives a fixed sound but this can be directed to output devices as you wish. For example this allows for having a master mix output and some other outputs for cueing or just partial mixes.

Although my main concern was to have a proper multithreaded mpg123 frontend that uses the EQ and makes gapless playback possible with working around the issues that the mp3 format and mpg123 have with that (update: new mpg123 can work around this for mp3s with lame tag), this is only part of it. There is the mixing engine with pitch (eh, speed, just speed) control and flexible setup and the possibility to add way less complex inputs and outputs.
Now there is a framework that lifts the need for any multithreading code from an input module author -- the thing will still be multithreaded also when you didn't write a line starting with pthread_ .

Additionally the multithreading makes _much_ sense concerning the network code: There is a thread for the main server watching the port and for every client connected. I really think that this is more appropriate than querying all sockets and the port in a loop from time to time (though I must admit that I didn't measure performance differences on that part... if you don't have many connections, the simple loop may still work well).

So, if you really don't care at all about the stuff that made me code this, then feel free to consider the "original" mixplayd. On the outside it should behave quite similar and may be just what you searched and its a smaller static binary (1/3 here). One part of the gapless story (called "autocue") has already went into mixplayd, too (version 0.60). But bear in mind: I have this thingy also on the end to catch mp3 frame / audio cd sector padding, not to forget what the "follow" command does.

In the end, I don't want to bash mixplayd. It inspired me to waste my study time with writing this ah-so-flexible mixing engine, and here it is;-)


Starting the daemon
===================

This says `dermixd -h`:

DerMixD
 - a tcp-controlled audio file mixing daemon in the tradition of mixplayd (well, on the outside) - 
Version 1.6.0,  Written and thus copyrighted in 2004-2008 by Thomas Orgis. Licensed under the GPL 2. No warranty of any kind. Have fun;-)

Usage: dermixd [-c|--console] [-r|--remote] [-h|--help] [other parameters]

Some configuration parameters can be given as "-o" for simple on/off options and "-p <value>" or "--param <value>" for parameters with values:
-c, --console	This disables daemon mode (stay in foreground, then).
-r, --remote	Allow connections from the outside. Default is to only accept connections from localhost. Consider that there is no authorization checking for clients, be it local or remote (so, at least  every user on your machine can talk to me)!
-p, --port	specify the TCP port to use
-n, --nothing	Start with no initial setup (otherwise: 2 inputs -> 1 default output)
-b, --buffer <positive_number>	specify mixer buffer size in samples
-a, --samplerate <positive_number>	specify mixer and output sAmple rAte
-o, --output <filename>	in daemon mode stdout/err get redirected to that file (debugging)
-O, --oldstyle	use the old - and deprecated! - practice of [say] and [watch] strings without command prefix
--version	just print name and version
--show-api	print out all control commands for use by clients (the API)
-h, --help	give this help info

And there are the more complete configuration parameters organized in groups you can specify as "group.name=<value>" on the command name or in the future via configuration file:

Parameter input: group.name=value; you can omit group "main"
"--" stops parameter parsing (for file names or similar following)
Example: program name=value main.another=anothervalue -- file1 file2

The parameters:
main.help:            give help	[no]
main.tcp:             listen on a TCP port	[yes]
main.unix:            listen on UNIX domain socket	[no]
main.socket:          UNIX domain socket path	[/dev/shm/dermixd.socket]
main.port:            specify the TCP port to use	[8888]
main.buffer:          mixer buffer size in samples	[2048]
main.audio_rate:      mixer and output sample rate	[44100]
main.channels:        channels for mixer and all outputs (1 for mono or 2 for stereo)	[2]
main.output:          in daemon mode stdout/err get redirected to that file	[/dev/null]
main.default_output:  default output device class to use in standard setup	[alsa]
main.default_outfile: default file for default output...	[default]
mpg123.decoder:       the decoder to start - ideally the path to some binary that behaves like mpg123-thor;-)	[mpg123]
mpg123.prebuffer:     mpg123 input prebuffer size in seconds (based on mixer samplerate)	[2]
mpg123.zeroscan:      on/off the facility to remove zero padding and beginning and end of tracks (helps gapless transitions)	[off]
mpg123.zerolevel:     specify an integer number for the raw level that is just not considered zero anymore when removing padding - you may want to avoid some dirt...	[100]
mpg123.zerorange:     specify an integer number of samples as limit for the scans	[4000]
mpg123.gapless:       start decoder with --gapless option (for mpg123 >= 0.60)	[off]
mpg123.nice:          start decoder via nice -n <adjustment>	[0]
alsa.buffer:          alsa buffer size in seconds	[0.1]
fifo.wait:            timeout for waiting for data on the decoder pipe in usecs; this could influence performance and will influence end-of-track triggering	[1000]


The control interface
=====================

The control interface is very similar to that of mixplayd. In fact, i made it use the same syntax for simple operations that my already coded frontend already used. But there's more (and less).

The communication consists of text sent through a tcp socket. The most easy way to open such a connection on your very own machine is

telnet localhost 8888

(8888 is the default port that mixplayd and dermixd use)

Then you type "something" and dermixd does something, giving a response (normally preceded by [something]) after that.
DerMixD uses UNIX line end for responses and accepts command lines ended by UNIX, DOS or MAC convention (line stopped by first \r or \n).

Specification of a command is the name of the command followed by its parameters -- in order -- divided by spaces. Commands accepting filenames/urls that may contain spaces itself take this as last parameter, so that with the fixed number of parameters there is no need for quotes or the like.

A command generally returns

[<command>] success

or

[<command>] success: <value>

when some single value was set (seeking counts as setting the position in seconds)

Also, there are commands that return some information in (possibly) multiple lines:

[<command>] +begin
first line
second line
...
[<command>] -end

An error is indicated by

[<command>] error: <some possible explanation / hint about what and why>

A client should be flexible with the whitespace, be it the number of which or the type (spaces, tabs), and not insist on a ":" after error/success . For example

[load]error
[load]         					  error:

Should be interpreted to mean the same thing. I'm not saying that dermixd will produce such extreme responses, but some small change in spacing for readability should not break existing parsers.

The actual form of <command> depends bit on what was issued. Normally it is the command name you used in the request, but special cases are script and say:

script 1 32 bass 1 0
[script/bass] success

The response of a script contains the scripted subaction name, if the _initial_ parsing was successful.

script 1 32 laod 
[script/load] error: parsing of subaction failed <- unknown command

 Otherwise:

script asf 
[script] error: unable to get (all) needed arguments

Then, the say command gives a 

[say] success

immediately and then

[saying] <the stuff it should say>

(You may say that that is without sense, but it is not when you use it in a script as notification.)


You can get the full listing of commands from `dermixd --show-api`:

Full API listing:

<name>:	<description> <flags if any>[; parameters: parm1(type) parm2(type) ...]
meaning of flags:
	container -- this action can contain a subaction (given as parameter); read: this is the script command
	nosub -- this action cannot be a subaction in a container; read: not allowed in script
	optarg -- parameters are optional: either specify all or none

volume:	set inchannel volume factor; parameters: channel(integer) volume(float)
bass:	set inchannel bass eq factor; parameters: channel(integer) value(float)
mid:	set inchannel mid eq factor; parameters: channel(integer) value(float)
treble:	set inchannel treble eq factor; parameters: channel(integer) value(float)
eq:	set all three eq factors; parameters: channel(integer) bass(float) mid(float) treble(float)
pause:	pause inchannel; parameters: channel(integer)
speed:	set playback speed factor (1=normal); parameters: channel(integer) value(float)
pitch:	increase/decrease playback speed factor; parameters: channel(integer) value(float)
stop:	stop playback on inchannel; parameters: channel(integer)
start:	start playback on inchannel; parameters: channel(integer)
seek:	absolute seek on inchannel; parameters: channel(integer) position in seconds(float)
rseek:	relative (from current position) seek on inchannel; parameters: channel(integer) offset in seconds(float)
bind:	bind input channel to output channel; parameters: inchannel(integer) outchannel(integer)
unbind:	release bond between input channel and output channel; parameters: inchannel(integer) outchannel(integer)
outpause:	pause outchannel; parameters: channel(integer)
outstop:	stop playback on outchannel; parameters: channel(integer)
outstart:	start playback on outchannel; parameters: channel(integer)
script:	script an action triggered by an inchannel passing certain time/position, command is just any valid (and allowed) command with arguments [ container nosub ]; parameters: channel(integer) time in seconds(float) command(string)
nscript:	script an action triggered by an inchannel passing certain time/position, script will be executed n times (one for each triggering) [ container nosub ]; parameters: channel(integer) n(integer) time in seconds(float) command(string)
showscript:	show the script commands for programmed actions on an inchannel [ nosub ]; parameters: the channel(integer)
delscript:	remove scripting actions from inchannel; parameters: channel(integer)
follow:	let one inchannel follow another (start folower gaplessy after leader stopped); parameters: leader(integer) follower(integer)
nofollow:	release the followership on a leader; parameters: leader(integer)
getstat:	give status info in mixplayd-like format (please do not use... may change/vanish completely) [ nosub ]
g:	short for getstat [ nosub ]
fullstat:	give full status info on all in- and outchannels (several lines) [ nosub ]
f:	short for fullstat [ nosub ]
say:	just say something (I put [saying] in front); useful when timed somehow; parameters: the message(string)
watch:	watch inchannel (get messages on events and ongoing playback); parameters: channel(integer)
unwatch:	stop watching inchannel (see watch); parameters: channel(integer)
play:	load track and start playback on inchannel [ nosub ]; parameters: channel(integer) track(string)
load:	load track on inchannel (be prepared for starting it) [ nosub ]; parameters: channel(integer) track(string)
inplay:	load track and start playback on inchannel [ nosub ]; parameters: channel(integer) driver(string) track(string)
inload:	load track on inchannel (be prepared for starting it) [ nosub ]; parameters: channel(integer) driver(string) track(string)
scan:	scan input properties [ nosub ]; parameters: channel(integer) list of properties(string)
outload:	load resource on outchannel with specified driver/device (prepared for activity) [ nosub ]; parameters: channel(integer) driver(string) resource(string)
outplay:	load resource on outchannel and start playback [ nosub ]; parameters: channel(integer) driver(string) resource(string)
length:	determine exact length of track (stops channel, decodes track till end, seeks back to where it left off); parameters: channel(integer)
preread:	read through a file once (to have it cached by file system/kernel) [ nosub ]; parameters: file(string)
addin:	add input channel (with optional nick name) [ nosub optarg ]; parameters: nick name(string)
remin:	remove input channel [ nosub ]; parameters: id(integer)
addout:	add output channel (with optional nick name) [ nosub optarg ]; parameters: nick name(string)
remout:	remove output channel [ nosub ]; parameters: id(integer)
id:	print my full identification [ nosub ]
close:	close current connection
fadeout:	dummy to remind you to do custom fading via script actions [ nosub optarg ]
vol:	set inchannel volume in percent; parameters: channel(integer) volume(float)
buffer:	load audio into mixer buffer (this includes zeroscan); parameters: channel(integer)
sleep:	let me sleep until any client wants something
shutdown:	lay down and die gracefully
help:	give some info / usage pattern on a command [ optarg ]; parameters: the command(string)
showapi:	show the whole API (list help for all commands) [ nosub ]
threadstat:	list spawned threads [ nosub ]
ls:	list file/directory [ nosub optarg ]; parameters: dir/file(string)
cd:	change directory (for client actions) [ nosub ]; parameters: dir(string)
pwd:	print current working directory [ nosub ]
spy:	spy on client communication [ nosub ]
unspy:	stopy spying on client communication [ nosub ]
showid:	toggle showing of channel id in responses to channel commands [ nosub ]; parameters: 1(on) or 0(off)(integer)
peer:	send a message to another peer (client). [ nosub ]; parameters: your peer name(string) peer's peer name(string) message(string)
addpeer:	add a peer entry, registering for messages [ nosub ]; parameters: peer name(string) description(string)
rempeer:	remove a peer entry [ nosub ]; parameters: peer name(string)
showpeers:	show a list of peers with optional description [ nosub ]


Scripting
=========

There are some special commands:

	script <ch>	<time> <command> <parameters>
	nscript <ch> <n> <time> <command> <parameters>

	showscript <ch>

	clearscript <ch>

These are used to manage a list of commands each to be executed when input channel <ch> reaches position <time> (see descriptions above for individual function).
Negative times have the special meaning of execution right after track end; the execution order of several scripts with negative times is according to the absolute value: The bigger the absolute value, the closer to the end (the most negative time is the first one).
Of course you can manipulate a different channel than the one giving the time; in fact, you're free to use all commands except the ones marked with the nosub flag in the full listing above.
DerMixD doesn't do fadeout on its own as mixplayd does - you have to draw the line somewhere... instead you can program your (cross)fading curve like

volume 1 0
script 0 40 start 1
script 1 0 volume 1 0.1
script 1 0.1 volume 1 0.2
script 1 0.2 volume 1 0.3
script 1 0.3 volume 1 0.4
script 1 0.4 volume 1 0.5
script 1 0.5 volume 1 0.6
script 1 0.6 volume 1 0.7
script 1 0.7 volume 1 0.8
script 1 0.8 volume 1 0.9
script 1 0.9 volume 1 1
script 1 1 volume 0 0.9
script 1 1.1 volume 0 0.7
script 1 1.2 volume 0 0.4
script 1 1.5 volume 0 0.2
script 1 1.8 volume 0 0.1
script 1 2 volume 0 0
script 1 2 stop 0

You have the power - use it!

The time argument has a special meaning when being smaller than zero: Then it means "do it after the track ended".
Also, smaller negative values put the action in front of others with bigger negative time values (in the mathematical sense of smaller and bigger):

script 0 -1 say you!
script 0 -2 say Hello 

will result in "Hello " being said first, then "you!".
The actions will all happen in the same time frame, but in the order specified by their begative times.


Available input devices
=======================

mpg123 - play mpeg audio files/urls (mp3, mp2) through mpg123 decoder
sndfile - play any file libsndfile can handle
vorbisfile - OGG/Vorbis through libvorbisfile
sine - generate sine tone, load with url scheme sine://<freq> or sine://<freq>@<sampling rate>
raw_s16 - raw stereo audio files, signed short (assumed to be 16 bits on your box! that should be done more straight in future), host byte order
dummy - could be called silence... because that is what it does, also good as basic template

The sndfile and vorbisfile inputs have to be enabled in compilation stage (via make VORBISFILE=yes / make SNDFILE=yes).

The raw input can be given parameters via the inload (and inplay) commands:

	inload 0 raw:1ch:48000Hz file.raw

Will play a mono file with 48000Hz appropriately.
Without any parameters the files are assumed to match the main mixer settings of mono/stereo and sample rate.


Available output devices
========================

There are some in various states of usefulness (stable not meaning rock-solid and foolproof but generally working without problems for me).

dermixd name (for outload)  description                                         state
oss[_threaded|_serial]      audio hardware output via OSS API,                  stable 
                            e.g. /dev/dsp; works with Alsa OSS emulation, too   
alsa[_threaded|_serial]     audio hardware output via Alsa API                  stable (better latency than OSS)
mme                         audio hardware output via MME API of Tru64 Unix     semi-stable**
                            (written for a Compaq XP1000)
text                        write ASCII text (TAB-separated numbers) to STDOUT  stable
raw_s16                     write RAW Signed 16Bit data to file                 stable

** not able to be _removed_ properly; I only tested this on a XP1000 box that plays fine except for some outage once or twice every minute. Could be some system service stopping the mmeserver, dunno.
One could use OSS on Tru64, though - perhaps that's better.

About threaded and serial modes:

The output can be run with a separate thread doing the output or with the main mixer thread doing this (serial mode). For a single output device you won't have a benefit from threaded mode but when having multiple outputs that is the way to do the waiting (while writing to device buffer, disk, ...) in parallel and prevent underruns because of doing one nothing after another.

The threaded mode is default now when loading oss or alsa; you can load a specific variant (regardless of default) with:

	outload 0 oss_serial /dev/dsp
or 	
	outload 0 oss_threaded /dev/dsp

The same goes for alsa_serial and alsa_threaded...


Version Policy
==============

I use a version scheme with three parts:

<major>.<minor>.<bugfix>

The major number stands for the basic feature set and interface. All 1.x versions shall (in a perfect world...) be compatible to any 1.y with y < x. Minor versions may add but not remove features and shall (...) not change default behaviour. This does not necessarily include the internal API for input/output modules since it has still to prove and possibly improve itself to actually host some more of them.

The last number is just for bugfixes that don't affect the nominal features (but may repair broken ones).

I don't discriminate even and odd numbers for stable/development releases or the like. If m > n, m should be better.

A client gets the interface version on connection init:

	shell$ telnet localhost 8888
	Trying 127.0.0.1...
	Connected to localhost.
	Escape character is '^]'.
	[connect] DerMixD v1.0

The full version specification including bugfix number and possible suffixes (-dev, -test2,...) is available via the id command:

	id
	[id] DerMixD v1.0.0

So, any client program that wants to use ensure that some functionality introduced in version 1.3 is available should parse the initial [connect] string and check for a minor version >= 3.
