Continuous Data Analysis with Analog Computers Using Statistical and Regression Techniques
GENERAL SECTION
COMPUTING TECHNIQUES: 1. 3. 2a
CONTINUOUSDATA ANALYSIS WITH ANAI,OG COMPUTERS USING
STATISTICAL AND REGRESSIONTECHNIQUES
ABSTRACT: This paper shows how certain fimdamental statistical
parameters of a given population--the mean, the variance, the autocorrelation frmction, the cross conelation function, the Fourier transform, or the power q)ectrum--can be estimated contlnuously througtr
calculations employlng simple analog techniques. tr addltlon, alalog
methods are developed for contlnuous regression analysis wherein two
or more populations or variables are compared statisticaUy to find
significant relationships.
*v
v
Prinr . d i n U . 5 . A . 2 6 4
O El .dr oni c As s o c i d r6 3 , I n . . 1 9 6 4
Al l R i s hr 3 R .!.N .d
Bol l .i i n N o. ALAC 6 2 0 2 3
CONTINUOUS DATA ANALYSIfI WITH ANALOG COMPUTERS USING
STATISTICAL AND REGRESSIONTECHNIQUES
t
GENERAL
The need for statistical data analysis intieprocess
industries is well known. Measurements ofprocess
variables or parameters are subject to random disturbances such as the presence of impurities in
varying amounts, environmental changes, weather,
etc, Very often it becomes necessary to obtain the
"best estimate" of a variableover somepriortime
interral for puryoses of control. It is this concept of
"estlmate" that intrcduces statistics.
With the availability of small, rugged and reliable
aJralog computing components especially designed
for plant environments, it becomes feasible both
economlcally and technically to apply statistical
techniques to the analysis of continuouB data for
either measurement or control. Of special importance is the fact that an alalog device can do a
simple or complex calculation task while remaining
a small package in terans ofphysica-I dimensions ard
cost. Other computational approeches almost aIways imply the purchase ofa relativelylarge "minimum" amount of hardware. Thus, one is encouraged to e:elore the 'lsimple" applications-- situa.tions which pay their own n'ay while prcviding experience in the use and testing of the analog approach.
Such computa.tions car be performed Off-Line, OnLine Open Loop, or On-Line Closed Loop on continuous-signal inputs; digitizing ofthe analog signal
is unnecessary. Notsy slgnal6, unavoidablein many
pilot or plant operations, whose deviation traces
serse as the basis for subsequent ca.lculations can
be "reduced" to more meaningful form by relative-
ly simple and economica.l computer circuits. The
mean and standard deviation of a noisy slgnal can
be recorded continuously arrd on-line, so that all
subsequentproblems in interpreting the data can be
reduced markedly. More complex but still relatively
inexpensive circuits can be used to record continuously, either on-line or off-line, the Fourier
Series Coefficients of a signal. Or, in the dynamic
testing of systems, the transformation from impulse response to frequency response can be accompllshed, thus permitting determination of the
best combination of simple input and easily-lnterpreted output,
THE MEAN
One of the filndamental statistical estimates isthat
of the t'mear" or the "arithmetic average" ofa
variable or a parameter. When dealing with discrete informatlon, the mean is defined by the summation
N
I
(1)
which is recognized easily as the familiar arithmetic average. For data analysis, this statistical
property is important for two reasons: 1) itis fundamental to the definitlon of other statistical parameters, and 2) lt applies equallywell to normal population distributions and to those that are not distributed normally.
It would appear desirable to utilize this statistical
property in the analysis of continuous data or for
the measurement of continuous process variables
for purposes of control. In order to do so. it becomes necessary to obtain the '.estimate', of the
mean as a continuous and changing function of
time. Slecifically, one must be able to define and
compute the average or mean value, d, of a signal,
f(t), varying with time over the interva.l Tf t_<T2.
As an example, assume that a steet mill is producing a continuous metal strip which, ideally,
should be uni.form but actually is fluctuating in
fhickness (as shown in Figure 1) because of inevitable random disturbances in the process. Over
+ f(u
YL
Figure 2, Anolog Circuil for Colculolion oI Estimote of lhe
Meon for o Fixed Time Infervol
A refinement to this circuit would be to Eenerate
I1t; as a continuously varying function of ti-me as
shown in Figure 3. The time interval then must be
considered as a variable so that the average is
computed continuously from time T,. In Figure 3,
T. is considered to be zero computerr time and T^
-t_z
has been replaced by t since the upper limit of the
integral is a variable.
ul
o
z
Yt'
T
;t,r={/',t,ro,
z
+ f 0)
F
LENO TH : J TI M E
F igure I. Th ickne ss o f St eel Plot e f r om o Rolling M ill o s o
Func t ion of Tim e
any time
interval
of
reasonable
length,
the mean
thiclnxess should be equal to the nominal value although small deviations are allowable. A sensing
device is monitoring the thiclaless as the strip
emerges from the mill, and a transducer is generating a signal, f(t), which is proportional to the
instantaneous thicl$ess. We would like to compute
the average value of this signal so that it car be
compared with the desired or nominal value in
order to see if the process is ulder control.
The most obvious definition of the mean or average value for f(t) over the interval T1<t<T2 is
Tz-Tt
f'
f(tl dt
t2\
'1
This value can be computed with the simple circuit
of Figure 2, At time T1 the integrator is placed
in the COMPUTE mode, arld at time T2 its output is
observed. The integrator then call be reset and
another average taken.
Figure 3. Anolog Circuit lor Colculotion ofo Continuous
Esfimole of the Meon for o Fixed Time lntervol
Y,
The circuit of Figure 3, although theoretically an
improvement over that of Figure 2, has two rather
obvious limitations: 1) the uncertainty of the division when t =0, and 2) the need to select maximum
rtmning time in advance, since the integrator will
eventually overload. This latter difficulty also occurs with the circuit of Figure 2. In each circuit,
the integration can continue only over a certain
range of time, and the circuit then must be reset.
However, the past values of I(t) are lost in the resetting; the averag€ computed during the second
!rrun" are independent of the va"luesoff(t) obtained
during the first run. If a succession of runs of
length T are made and the circuits are reset each
time, it is clear that the last computed averag€
depends only on the behavior of f(t) in the last T
units of time. !1 other words, fmm the point of
view of the most recent average, information older
than T units of time is obsolete.
The resetting necessary with the previous two circuits ca-n be avoided, and a much simpler circuit
obtained without worry about overloads and division circuits, The clue to the method is the fact
that past values of I(t) become obsolete. Since the
\a,
basic signal, f(t), ls continuous, it seems advantageous to let past infornation become obsolete
graduauy rather than abrrptly. This means that
f(t) has to be defined in such a way that recent
values court much more heavily tha.nearlier va.lues
a^ndtlte behavior of f(t) in the remotepast has very
little effect. This suggests that a weighted average
be used.
This
The weighted-average-i(t)
of a function, f(t), over
T1 S t < T, with weight function C (t) is defined by
($+ wher6 d (t))0 in the interval T1<IST2. Thus
The minus infinity in the lower limit serves to indicate that the average has been generated for such
a long time that the effect of what happened before
T1 is negligible. In otler words, since the exponential weighting function, e*', approaches zero as
t* -co, the importance ofevents prior to T1 is negllgtble if T1 is suitably chosen.
T^
T,'
f(t) o(t) dt
f( t) =
(3)
T^
T"
c (t) dt
can be simplified
T^
fa
i1T; = o.-.'T,
|
f1r1
' =o
f(t) =
f:
I
eot f1q dt
(4)
-2
d,L
edt
T
f (T) = a
(7)
- t) 61
r1t1
"-"(T
(s)
I
Otterana.n (2) defines this to be the '. Exlponentiauy
Mapped Past" or EMP of f(t) over a time interva.l
defined by d '.
Implementation of the analog circuitfor solvingthis
equation is reasonably straightforward. Differentiating Equation 7 with respect to machine time,
T, (t is a duIIIIny variable) gives
..
T
-aT
= a {.(--ce
,["
.J
i
T^
fz
l^r1
o tz
e
-@
.ot f1t1dt
ot
rlty at
. ^-qT
(e)
t""trtotf
(5)
f (t) = d - - - - E- .- ..;i-
o t1
-e
*Numbe's in porentheses in the body o( the iext reter to rcfcrcnces
I isl e d i n A P P E N D I X I,
r6y
at
I
"-oT J_"ot
1
-
T
Re- arranging
Remembering the requirement that the recent past
must be emphasized and the remote past de-emphaslzed, it follows that we should chooseaweighting function, C (t), which is increasing and such that
Iim d (t) = 0' Many functions have this property
but
the exponential function is a natura"l one and leads
to a eimple computer circuit. Pickingan exponential
weighting function, eat (d>0), Equation 3 becomes
(6)
r1t1at
Dropptng the subscripts, Equation 6 car be written
as
f
The integral in the denominator serves to .,normalize,' the e4ression. The function O (t) can be
chosen arbitrarily to emphasize or de-emphasize
various parts of the interval from T1 to T2.
""t
J_6
Tt
o
by letting Tr+-co, or
d f (r)
dT
=
"[-i(1)
+ r(r)] = arlq -af 1q (10)
tThose tonilio. wirh lineot onalysis dnd, in porticvtoL convorurion
int.sftls, wi ll rccognlze Equorion 8 os re ourpur ot o firre, whose
inpulse response is aeat; thot is, o ,trsr-arder litrer wirh ti||,.
Equatlon 10 is lmplemented by the simple chcult
of Figure 4, which is recognizedeasily as the circuit for a simple filter or firetorderlag. Note that
+t(tl
a
-t0t
In the circult of Flgure 4 it is obvlous that an inltlal
condition applied to the integrator wiU improve the
computed average at the b€ginning. Thts value should
represent a goodguess asto the nominal or er<pected
mean value of f(t). Onenornallywould have such an
estimate avallable. Ifit is a goodestimate, the computed average will be reasonable from the start; if lt
is a bad one, lt wlll not mal<eany dlfference after
about three to five time constants.
V/.
THE VARIANCE
Figurc4. AnologCircuit6r Obtoining
rhe EMpEstimote
o{
theMeon
the input and output signals have been written in
terns of the more famiuar notation for time, t,
which is not to be confused with the dummv variable
of Equation 7.
The value of the constant, d, determines how fast
pest inlormEtion becomes obsolete. It is chosen
arbitrarily to be large enough to filter out nonessential random fluctuations and small enough
not to obscure long term trends. A useful rule of
thumb can be developed by examining the response
of the circuit_of Figure 4. ff f(t) chang€s abruptly
(step input), f(t) wiU fouow g"aduaUy, making 95%
of the change 1n 3 time constants or a time interval
of 3/q.In other words, as shown in Figure 5, after
!
A second important statistical parameter is the
varirurce which is used to give a basic measure of
the distribution of a population. It is defined as the
square of the standard deviatlon and is equal to the
mean-squared deviation of the variable from its
mean. For discretedata, an estimateof the variance
is obtained vrith the summation
s
2
r. ; 12
tt-')
1 L
N-
i=t
_-i.
v
(11)
The term N - 1 corresponds to the number of
degrees of freedom lnvolved in the calculation of
the estimates of the variance (3); practical considerations dictate that the number of samples,
N, will be larger than one.
W
FUNCTIOI{TO BE
d
VERAGED
lrl
o
In a marurer similiar to the definitlon of the mean,
Otterman (2) defines the EMP variance as
a
TEIOHTII{G
FUiCTtOil
c.
.I
I
,'g1 = .l
J _@-
I _---)
Figure 5. The EMP Meon o{ o Conlinuous Vorioblc Provides
o M.o"ut. of the Averoge o{ the Vorioblo for o Continuously
Updored Fixcd Tirns lniervol. Noterhs 95% dccreose in lhe
voluc of lhe weighting function over o pcriod of lcngth 3'la'
This meons thot th. wlighled ovcrogc ol time, l. i5 vittuolly
indepcndent of voltes thot occurrcd prior to lime t - 3/a'
three time constarrts, the integrator has forgotten
95% of tbe informatlon itbad before the step change.
Consequently, the EMP averag€ defined by Equation
8 is an estimate* of the mean over a time interval
approximately equal to 3/a.
*lt
a 99% .it
''ior.ly 5/a.
tion wcre used, tAe ti'..
int.Nol
would be opptcxi-
(r - t) dt (12)
frtu - rttll' er
-
Y"
1t
which, based on the preceding development of the
EMP mean, wlll be recognized as the welghted
average of the square of the deviation of the
variable from lts mean,
The computer circuit for calculating an estimate of
the EMP variance is developed easily without recourse to mathematical manipulatlons. From the
definition, and remembering that averaging is accomplished by the first-order filter circuit, the
following operations are requlred:
1) form the mean with the first-order
circuit.
ftlter
Y'
2) subtract the mean from the current value
of f(t).
v
3) square the difference of the mean from the
current va.Iueof f(t).
4) average the square with a second filter
circuit.
From these requirements the circuit of Figure 6
is derived easily,
r)-i(rt
.,-U
Figure 6. Anolog Circuit for Colculotior of rhe EMP Esrimote
o{ t he Vor ionc e
Example: The LD Steel Process car be used as an
example of the use of the EMP mearr and variance
for the control of a process. Thls is the oxygen
steel making process wherein it is possible to control bath temperatures without an external fuel
supply by charging the vessel with materials that
are thermally balanced. The charge materials consist of hot metal (lron), scrap, and lime,
The hot metal temperature can range from 2200'F
to 2600'F, and, hence, it is necessary to measure
the temperature of the iron to obtain a correct
thermal balance.
A two-color radiation pyrometer method is usedto
measure the iron temperature while it is being
poured into the vessel. A typical trace is shown in
Figure 7. (For further details of the process the
reader is referred to reference 4.)
2 600
At present, the "temperature" reading is inserted
manually into the charge balance computer (the
mea.n value of temperature is "guesstimated" by
the operator). This could be automated easily with
an EMP mear value circuit since the transducer
signal is a continuous electrical signal. The condition that the reading of the mean value circuit
should not be used at the treginning of tlle time
history, t<t1, for reasons mentioned previously,
can be automated by using the standard deviation
(variance) as a control criterlon, i.e., when a21T1
rj greater than a reference va.lue, do not use
f(T), when oz(T) is less than a reference value,
use f(t). The reference value chosen witl depend
on the maximum variance expected during smooth
pour conditions. This can be mechanized readily
on the analog computer by means of a comparator.
with this simple technique a better estimate of the
mean temperature could be inserted into the charge
computer automatically ard economically.
AUTOCORRELATION
The autocorrelation function, defined as an integral
between fixed limits, is converted easily to a continuous EMP autocorrelation frmction, O (T), by the
definition
T
f(t) f(t -r) e-d(r'
o(r) ='l
J__
- t) dt
2400
2300
lO
ll
(13)
Cross correlation a.1socould be accomplished by
the substitution of a second function, g(t), into the
time delay box, r, shown in Figure 8, so that the
output of the delay box is g(t - i) and the output of
the multiplier becomes - [rttt] [eA .
')]
READINGTAKEN
2 500
these readings until the smoke has b€en blorn
away (by a fan) and smooth pouring is established,
Note that even after these conditions have been
attained, the temperature measurement, tft<t2,
is subject to fluctuations.
t2
TI M E
Figure 7. Plol of Hol Melol Tempcroiur€vs Time lor fhe LD
Steel Process
TIME
DELAY
[ttr][t-.r]
\y
The initial variations in temperature, t6(t(11,
are due to the presence of smoke ard the formation of voids in the pour, The operator disregards
Figure8, Circuit {or Obroiningthe ContinuousEMP
Aulocorrelolion
Funclion6r TirneDeloyr
Reasonable time delays are obtained easily by assembling linear ana-logcomputing components. Figure 9 shows a fourth-order Paddcircuit for g€nera-
F(o). The real component, E1, is formed from the
transfer fuction
F
-L
( P +d ) d
T(|\-------.,-----T
.\!/
+
N O T E .715r =O . 4 6 6 7 (1 ,/r)
(p
ra)-
(18)
vi
c,
and the imaginary component, E2, from
F.
"2
f(t)
ol an
(1e)
1p + oyz + .2
t/t
Figure 9. Circuil for Fourth-OrderPode Approximotion{or
ldeol Time Deloy of Mogniruder (5)
ting atimedelay, r. This circuit is accurate to within 1 degree of phase shift for input frequencies in
f(t) such that the product of the maximum useful
signal frequency, o*, with the time delay, r, shall
not exceed 6,5 radians. i.e,.
< 6.5 radians
7,.,
-m-
Note that this circuit gives P(r,r) at one value of (,.
The parameter, o, can be changed merely by
changing the two potentiometers labelled "@".
YrJ
(14)
FOIJRIER AND POWER SPECTRT'M ANALYSIS
The EMP Fourier transform of f(t) is defined as
Figure 10. Circuit for Colc,rlolion of EMP Fouder Tronsform
ond Power Spectrum
T
F((,) = dl f(t) e
J_*
or
-a(T - t)
-iot ..
eda
(15)
(16)
T
F1<.r;= o
"-jcoT
I-
f(t) e-a(T
- t)
- t)
u,
"jc"(T
An alternative is to build similar circuits in
parallel, all having the same input, f(t), and differing only in the setting of <.r. This will a1low
many points of the power spectrum to be obtained
simultaneouslyREGRESSIONANALYSIS
The EMP power spectrum i8 defined as
(17)
P((') =
Y,,
l"t'lI '
rv--,n
- t)
"-o(T
cos <.: (T - t dt]
"'V-n"
-c(T- t) si"r(T- t) dtl
Figure 10 shows the analog circuit for obtainingthe
power spectrurn, P(o), of the Fourier traasform,
'L
Until now, only those statistical parameters that
describe a single population--mean, variance,
power spectrum, etc.--have been discussed. Of
interest, also, is the method of statistics whereby
relationships between two 01: more populations,
representing different variables, are found. This
method is called "regression analysis".
There are many types of regression---linear,
quadratic, high order, multivariate, etc. These
terms refen to the type of expression used to relate the variables. For example,
v
y = rnx + b
(Iinear regression)
y = ax
(quadratic regression)
+ bx2 + c
(20)
(21)
\r/
Regression consists, essentially, of finding the
to a set of data using a least-squares
criterion. While several authors have hinted at
obtaining a "least squares" fit by analogtechniques
for special cases, none have shown a straightforward solution to the regression problem as defined above for continuous variables.
Figure 11 shows the analog computer circuit for
calculating the "least squares" parameters ,?
and b using the definite integral. One should note
As an illustration of how least squares fitting
would be performed on the analog computer, consider the linear case defined by Equation 20, For
this general equation it can be shown (6) that the
following two equations will define the unlqrown
par:rmeters rn and b, the slope and intercept of
the line, respectively.
NN
lv,=-lx ,* N u
Lr
NN. N
lxv. rt = m \/-t
i=l
(22)
I
LJ
/J
i=l
xi1 * \-
/-t
i=l
x-I
(23)
Since X and Y will be continuous fr.mctions of
time. the discrete summation from l to N can
be replaced by a time integral where the total
time, t, is proportional to N. Therefore, Equations 22 and 23 become
I. tt * =-;ft***r,
(24).
tt
I o o r= ^[
Jo
x2 d t+r/xa t
(25't
Jo
These two equations can be solved simultaneously
to yield rn and b as follows:
o-It"
* - ^It* *
(26)
v
(27)
9l l
-xz
ra-
Joxdt
the"LeoslSquores"
Figurell. Circuit6r Obtoining
Regression
Poromelers
m ondb
that in calculating ,? and b from the definite integral we form an r'algebraic" Ioop. This brings
up the question of circuit stability. It can be shown"
that once the computation is under way, (t>0), the
loop will be stable unless X is constant. However,
if X is constant it cannot be used as an independent
variable in a correlation study. Overloads at the
very beginning of the computation (a region of no
interest) can be talen care ofwith feedback limiters
on the division amplifiers, or by using a "steepest
descent" division circuit (7).
It should be observed that ,x and b are defined at
every instant of tlme. For small values of t
(conesponding to small sample size), the estimates
of rn and b wiu be relatively insignificant and,
hence, will be changing rapidly. As the time interval increases. however. the values of n a-ndb
become more significant and actually should reach
"steady state' or non-changing values.
The technique for linear regression can be extended
to quadratic or higher order regressions, it then
being necessary to define a new set of equations-such as Equations 22 and 23-- fo]r determining the
unknown pararneters of the regression system.
Once the equations are defined, they can be converted to continuous integrals and instrumented by
standard analog techniques.
Again we are dealing wlth continuous signals, which
means that there must be a limit of the integration
interva-l if circuits such as that of FiEure 11 are to
*see Reference (8J
b€ used. Just as before, the need for resetting of
integTators can b€ eliminated by converting the
equations for rn and b to EMP equations and, thereby, obtaining truly continuous estimates of the regression parameters.
Equations 22 and 23 are rewritten
through N. This yields
v
by dividing
N
N
Fx.
Sv. L
Z-/
L/ L
N
--- N
(28)
I *,', \- xi
N
N.
i=1
i=1
N
N
N
\-
x.
/Jl
+ h -:-:"N
(2e)
Figure 12. Unscoled Anolog Circuif for Colculotion o{
Continuous EMP Yolues of Regression
Porcmelersm ond !
v
CONCLUSIONS
Recalling the correspondence between EMP variables and discrete summations, one ca.ntransform
Equations 28 and 29 immediately into continuous
EMP notation which gives
Y= m X
+b
F=m X2+b X
(30)
(31)
where m and b are no\r' the EMP estimates of the
regression parameters. The analog circuit required
is shown in Figure 12. The €tatements made with
regard to the stability of the circuit shown in Figure 11 apply also to the algebraic loop formd in the
Figure 12 circuit,
It should be observed that time hae, in effect, been
talen out of the problem by the conversion to EMP
variables. There is no longer any need to ..reset"
the integ?ators since they are now serving as convolution circuits rather than pure a.ccumulators. It
follows that circuits similiar to those showncan be
instnrmented for continuous higher order arrd continuous multi-variable regressions. All that is re.
quired is more analog computing equipment.
The conversion of statistical parameters to EMP
variables enables data analysis to be performed
continuously through the use of relatively simple
analog circuits. Limits on the size of the integration interval norma-lly encountered with continuous
signals have been eliminated; the need to ,.reset"
the integrators is no longer required since they are
now serving as convolution circuits rather thanpure
accumulators. A continuous estimate of statistical
parameters can be cafculated readily.
The concept of replacing discrete summationswith
the EMP mean can be a va-luableone. h addition to
obvious uses for instrumentation and control, for
both on-line and off-line systems, this technique
a-lso can be used in analog simulation studies, For
o(ample, it is sometimes deslrable to calculate
the rms value of a computed variable. This is accomplished quite simply by 1) squaring the instantareous value of the variable, 2) taking the
mean of the square of the variable with an EMP
circuit, and 3) taking the square root of the mean.
Other uses arise in simulationworkwhere Gaussian
noise is used to disturb a particular parameter.
Y'
tl
v
APPENDD<I: REFERENCES
(1) Davenport, W.B., Jr.,
arrd W.L. Root: '.A! Introduction to the Theory ol Random Signals
and Noise", McGraw-HiU noot Company,ffi
(2't Ottermar,
Joseph: "The Properties
end Mefud
Mapped Past Statistlcal Valiables", IRE TRANS. on Automatic Control, Volume AC-s
Number 1, January 1960, pp. 11-17.
(3) Vol\ William: "Applied Stattstics fo. Engineers", Mccraw-Hill
New York, 1958. p. 136.
(4) Slato€ky, W.J.: t'End Point Ternperature Control
Metals, Volume 12, March 19€0, pp. 226-230.
Book Company, lnc.,
in LD Steel Making", Jouma-l of
(5) Brenner, M.M., and J.D. Kermedy: "Dead Time Slmdation for Electr:onlc Analog ComNational Sirnulation Council, December:11, 1957,
flE-:,,
(6) Widder. D.V.: Advarced Calculus. Prentice-Hall. 194?. P. 108.
(7') Favreau, R.R,! 'rDividing Circuit
Princeton Comp
New Jersey.
Obtained by Applying qgthod of Sleepest Ascent",
(8) HaDnaner, George: "Algebraic Loop8 - Some Stability CoEsideratiolEt' Educatlon and
Trairilg Memo #22, Electronic Aasociate€, Itrc., PrincetoE, New Jerael'.
v