UNCLASSI FIED/ / FOR OmCIAL USE ONLY Defense Intelligence Reference Document J^^^^^^ff Acquisition Threat Support 11 March 2010 ICOD: 1 December 2009 An Introduction to the Statistical Drake Equation UNCLASSIFIED/ /FOR QEEICTftl IISFANIY UNCLASSIFIED//POR OFFICIAL USE ONh¥ An Introduction to the Statistical Drake Equation Prepared by: Acquisition Support Division (DWO-3) Defense Warning Office Directorate for Analysis Defense Intelligence Agency Author: AAP Person 80 A dm inistrative N ote COPYRIGHT WA RNING: Further dissemination of the photographs in this publication is not authorized. This product is one in a series of advanced technology reports produced in FY 2009 under the Defense Intelligence A gency, Defense Warning Office's A dvanced A erospace Weapon System A pplications (A A WSA ) Program. Comments or questions pertaining to this document should be addressed to |A A P Person 1 | A A WSA Program Manager, Defense Intelligence A gency, A TTN: CLA R/DWO-3 , Bldg 6000, Washington, DC 203 40-5100. ii UNCLASSI FI ED/ /fOR OFFICIAL USE ONLY U NCLASSI FIED/ / POR OmCIAL UG E ONLY Contents 1. Introduction............................ ..iv 2. The K ey Question: How Far are They ?...................................... 4 3. Computing N B y Virtue of the Drake Equation (1961).......... 7 4. The Drake Equation is Over-Simplified................................. 10 5. The Statistical Drake Equation.............................................................................11 6. Solving the Statistical Drake Equation B y Virtue of the Central Limit Theorem (CLT) of Statistics............................................................................................13 7. An Example Explaining the Statistical Drake Equation.......................................15 8. Finding the Probability Distribution of the Et-Distance B y Virtue of the Statistical Drake Equation................................. 18 9. The "Data Enrichment Principle" as the B est CLT Consequence Upon the Statistical Drake Equation (Any Number of Factors Allowed)........................23 10. Conclusions........................................................................ 23 Appendix A: Proof of Shannon's 1948 Theorem Stating That the Uniform Distribution is the "M ost Uncertain" One Over a Finite Range of Values............................................................................25 Appendix B : Original Text of the Author's Paper #IAC-08-A4.1.4 Entitled the Statistical Drake Equation.............................................28 References................................................................................................................55 iii UNCLASSIFIED//JOJUamCIAl II^FfiNIY UNCLASSI FIED/ / FOR OmCIAL USE ONLY An Introduction to the Statistical Drake Equation 1. Introduction SETI (an acronym for "Search for Extraterrestrial Intelligence") is a relatively new branch of scientific research, having begun only in 1959. Its goal is to ascertain whether alien civilizations exist in the universe, how far from us they exist, and possibly how much more advanced than us they may be. As of 2009, the only physical tools we know that could help us get in touch with aliens are the electromagnetic waves an alien civilization could emit and we could detect. This forces us to use the largest radiotelescopes on Earth for SETI research, because the higher our collecting area of electromagnetic radiation is, the higher our sensitivity is (that is, the farther in space we can probe). Yet, even by using the largest radiotelescopes on Earth (the 310-meter dish at Arecibo, for instance), we cannot search for aliens beyond, say, a few hundred light years away. This is a very, very small amount of space around us within our galaxy, the M ilky Way, that is about 100,000 light years in diameter. Thus, current SETI can cover only a very tiny fraction of the galaxy, and it is not surprising that in the past 50 years of SETI searches, NO extraterrestrial civilization was discovered. Quite simply, we did not get far enough! This demands the construction of much more powerful and radically new radiotelescopes. Rather than big and heavy metal dishes, whose mechanical problems hamper SETI research too much, we are now turning to "software radiotelescopes," where a large number of small dishes (ATA = Allen Telescope Array, and ALM A = Atacama Large M illimeter/submillimeter Array) or even just of simple dipoles (LOFAR = Low Frequency Array) using state-of- the-art electronics and very-high-speed computing can outperform the classical radiotelescopes in many regards. The final dream in this field is the SK A (= Square K ilometer Array), currently being designed and expected to be completed around 2020. 2. The K ey Question: How Far are They ? But still, the key question remains: how far are they? O r, more correctly, how far do we expect the NEAREST extraterrestrial civilization to be from the Solar System in the galaxy? This question was first faced in a scientific manner back in 1961 by the same scientist who also was the first experimental SETI radio astronomer ever: the American, Frank Donald Drake (born 1930). He first considered the shape and size of the galaxy where we are living: the Milky W ay. This is a spiral galaxy measuring some 100,000 light years in diameter and some 16,000 light years in thickness of the Galactic Disk at half­ way from its center. That is: The diameter of the galaxy is (about) 100,000 light years, (abbreviated ly) i.e., its radius, RGaiaxy, is about 50,000 ly. iv UNCLASSIFIED/^f OR OmCIAh USE ONLY U N CLASSI FI E D/ / FOR OFFICIAL USE ONLY The thickness of the Galactic Disk at half-way from its center, hGalan, is about 16,000 ly. The volume of the galaxy may then be approximated as the volume of the corresponding cylinder, i.e. ^G alaxy — ^ ^G alaxy ^ • ( 1) Now consider the sphere around us having a radius r. The volume of such a sphere is _4 f ET_Distance 'Our _Sphere 2 In the last equation, we had to divide the distance "ET_Distance" between ourselves and the nearest ET civilization by 2 because we are now going to make the unwarranted assumption that all ET civilizations are equally spaced from each other in the galaxy! This is a crazy assumption, clearly, and should be replaced by more scientifically-grounded assumptions as soon as we know more about our Galactic Neighborhood. At the moment, however, this is the best guess that we can make, and so we shall take it for granted, although we are aware that this is a weak point in the reasoning. Furthermore, let us denote by N the total number of civilizations now living in the galaxy, including ourselves. O f course, this number N is unknown. W e only know that W >1 since one civilization does at least exist! Having thus assumed that ET civilizations are UNIFO RMLY SPAC ED IN THE GALAXY, we can then write down the proportion: ^G alaxy ^Our Sphere - n _ f (2) That is, upon replacing both (1) and (2) into (3): 4 (ET Distance — 77 I ----------------------- ^ ^G ctidxy^ 1 (4) The last equation contains two unknowns: N and ET_Distance, and so we don't know which one it is better to solve for. However, we may suppose that, by resorting to the (rather uncertain) knowledge that we have about the Evolution of the galaxy through the last 10 billion years or so, we might somehow compute an approximate value for N. Then, we may solve (4) for ET_Distance thus obtaining the (AV ERAGE) DISTANC E BETW EEN ANY PAIR O F NEIGHBO RING C IV ILIZATIO NS IN THE GALAXY (DISTANC E LAW ) 5 U N CLASSI FI E D//FOR OFFICIAL USE OM I Y U N CLASSI FI E D/ / FOR OFFICIAL USE ONLY ET_Distance(N) = (5) where the positive constant C is defined by C = ^6^~^^«28845 light years. (6) Equations (5) and (6) are the starting point to understand the origin of the Drake equation that we discuss in detail in Section 3 of this paper. Let us just complete this section by pointing out three different numerical cases of the distance law (5): • W e know that we exist, so N may not be smaller than 1, i.e., w > 1. Suppose then that we are alone in the galaxy, i.e., that /V =l. Then the distance law (5) yields as distance to the nearest civilization from us just the constant C , i.e., 28,845 light years. This is about the distance in between ourselves and the center of the galaxy (i.e. the Galactic Bulge). Thus, this result seems to suggest that, if we do not find any extraterrestrial civilization around us in these outskirts of the galaxy where we live, we should look around the Galactic C enter first. And this is indeed what is happening, i.e., many SETI searches are actually pointing the antennas towards the Galactic C enter, looking for beacons (see, for instance ref. [1]). • Suppose next that /V =1000, i.e. there are about a thousand extraterrestrial communicating civilizations in the whole galaxy right now. Then the distance law (5) yields an average distance of 2,885 light years. This is a distance that most radiotelescopes in Earth may not reach for SETI searches right now: hence the need to build larger radiotelescopes, like ALMA, LO FAR, and the SKA. • Suppose finally that /V =1000000, i.e., there are a million communicating civilizations now in the galaxy. Then the distance law (5) yields an average distance of 288 light years. This is within the (upper) range of distances that our current radiotelescopes may reach for SETI searches, and that justifies all SETI searches that have been done so far in the first fifty years of SETI (1960-2010). In conclusion, interpolating the above three special cases of N, we may say that the distance law (5) yields the following key diagram of the average ET distance vs. the assumed number of communicating civilizations, N, in the galaxy right now (Figure 1): 6 UNCLASSI FIED/ >EQR QEEICT M IISFOhllY U N CLASSI FI E D/ / FOR OFFICIAL USE ONLY Average DISTANCEof the nearest ET civilization vs. the ASSUMED NUMBER of ET civilizations in the Gali Figure 1. DISTANCE LAW; i.e., the Average Distance (plot along the vertical axis in light years) Versus the NUM B ER of Communicating Civilizations ASSUM ED to Exist in the G alaxy Right Now 3. Computing N B y Virtue of the Drake Equation (1961) In the previous section, the problem of finding how close the nearest ET civilization may be was "solved" by reducing it to the computation of N, the total number of extraterrestrial civilizations now existing in this galaxy. In this section the famous Drake equation is described, that was proposed back in 1961 by Frank Donald Drake (born 1930) to estimate the numerical value of N. W e believe that no better introductory description of the Drake equations exists other than the one given by C arl Sagan in his 1983 book "C osmos" (ref. [2]), in its turn based on the famous TV series "C osmos." So, in this paragraph we report C arl Sagan's description of the Drake equation unabridged. "But is there anyone out there to talk to? W ith a third or a half a trillion stars in our Milky W ay galaxy alone, could ours be the only one accompanied by an inhabited planet? How much more likely it is that technical civilizations are a cosmic commonplace, that the galaxy is pulsing and humming with advanced societies, and, therefore, that the nearest such culture is not so very far away - perhaps transmitting from antennas established on a planet of a naked-eye star just next door. Perhaps when we look up at the sky at night, near one of those faint pinpoints of light is a world on which someone quite different from us is then glancing idly at a star we call the Sun and entertaining, for just a moment, an outrageous speculation. 7 UNCLASSI FI ED//rOR OFFICIAh USE ONLY U N CLASSI FI E D/ / FOR OFFICIAL USE ONLY It is very hard to be sure. There may be several impediments to the evolution of a technical civilization. Planets may be rarer than we think. Perhaps the origin of life is not so easy as our laboratory experiments suggest. Perhaps the evolution of advanced life forms is improbable. O r it may be that complex life forms evolve more readily, but intelligence and technical societies require an unlikely set of coincidences - just as the evolution of the human species depended on the demise of the dinosaurs and the ice­ age recession of the forests in whose trees our ancestors screeched and dimly wondered. O r perhaps civilizations arise repeatedly, inexorably, on innumerable planets in the Milky W ay, but are generally unstable; so all but a tiny fraction are unable to survive their technology and succumb to greed and ignorance, pollution and nuclear war. It is possible to explore this great issue further and make a crude estimate of N, the number of advanced civilizations in the galaxy. W e define an advanced civilization as one capable of radio astronomy. This is, of course, a parochial if essential definition. There may be countless worlds on which the inhabitants are accomplished linguists or superb poets but indifferent radio astronomers. W e will not hear from them. N can be written as the product or multiplication of a number of factors, each a kind of filter, every one of which must be sizable for there to be a large number of civilizations: • Ns, the number of stars in the Milky W ay galaxy. • fp, the fraction of stars that have planetary systems. • ne, the number of planets in a given system that are ecologically suitable for life. • fl, the fraction of otherwise suitable planets on which life actually arises. • fi, the fraction of inhabited planets on which an intelligent form of life evolves. • fc, the fraction of planets inhabited by intelligent beings on which a communicative technical civilization develops. • fL, the fraction of planetary lifetime graced by a technical civilization. W ritten out, the equation reads N = Ns-fi> ■ ne-fl ■ fi ■ fc ■ fL (7) All of the fs are fractions, having values between 0 and 1; they will pare down the large value of Ns. To derive N we must estimate each of these quantities. W e know a fair amount about the early factors in the equation, the number of stars and planetary systems. W e know very little about the later factors, concerning the evolution of intelligence or the lifetime of technical societies. In these cases our estimates will be little better than guesses. I invite you, if you disagree with my estimates below, make your own choices and see what implications your alternative suggestions have for the number of advanced civilizations in the galaxy. O ne of the great virtues of this equation, due to Frank Drake of C ornell, is that it involves subjects ranging from stellar and planetary astronomy to organic chemistry, evolutionary biology, history, politics and abnormal psychology. Much of the C osmos is in the span of the Drake equation. 8 UNCLASSI FIED//FOR OFFICIAL USE OM I Y U N CLASSI FI E D/ / FOR OFFICIAL USE ONLY W e know Ns, the number of stars in the Milky W ay galaxy, fairly well, by careful counts of stars in a small but representative region of the sky. It is a few hundred billion; some recent estimates place it at 4 x 1011. V ery few of these stars are of the massive short­ lived variety that squander their reserves of thermonuclear fuel. The great majority have lifetimes of billions or more years in which they are shining stably, providing a suitable energy source for the energy and evolution of life on nearby planets. There is evidence that planets are a frequent accompaniment of star formation: in the satellite systems of Jupiter, Saturn and Uranus, which are like miniature solar systems; in theories of the origin of the planets; in studies of double stars; in observations of accretion disks around stars; and is some preliminary investigations of gravitational perturbations of nearby stars.1 Many, perhaps even most, stars may have planets. W e take the fraction of stars that have planets, fp, as roughly equal to 1/3. Then the total number of planetary systems in the galaxy would be Ns fp ~ 1.3 x 1011 (the symbol ~ means "approximately equal to"). If each system were to have about ten planets, as ours does, the total number of worlds in the galaxy would be more than a trillion, a vast arena for the cosmic drama. 1 Carl Sagan was writings these lines back in the 1970's, when no extrasolar planets had been discovered yet. The first such discovery occurred in 1995, when Michel Mayor and Didier Queioz, working at the "Observatoire de Haute Provence" in France, discovered the first extrasolar planet orbiting the nearby star 51 Peg. This first extrasolar planet was hence named 51 Peg B. Many more extrasolar planets were discovered around nearby stars ever since. As of April 2009, 347 extrasolar planets (exoplanets) are listed in the Extrasolar Planets Encyclopaedia. In our own solar system there are several bodies that may be suitable for life of some sort: the Earth certainly, and perhaps Mars, Titan and Jupiter. O nce life originates, it tends to be very adaptable and tenacious. There must be many different environments suitable for life in a given planetary system. But conservatively we choose ne-2. Then the number of planets in the galaxy suitable for life becomes Ns fp ne ~ 3 x 1011. Experiments show that under the most common cosmic conditions the molecular basis of life is readily made, the building blocks of molecules able to make copies of themselves. W e are now on less certain grounds; there may, for example, be impediments in the evolution of the genetic code, although I think this is unlikely over billions of years of primeval chemistry. W e choose fl ~ 1/3, implying a total number of planets in the Milky W ay on which life has arisen at least once as Ns fp ne fl ~ lx 1011, a hundred billion inhabited worlds. That in itself is a remarkable conclusion. But we are not yet finished. The choices of fi and fc are more difficult. O n the one hand, many individually unlikely steps had to occur in biological evolution and human history for our present intelligence and technology to develop. O n the other hand, there must be quite different pathways to an advanced civilization of specified capabilities. C onsidering the apparent difficulty in the evolution of large organisms, represented by the C ambrian explosion, let us choose fi x fc = 1/100, meaning that only 1 per cent of planets on which life arises actually produce a technical civilization. This estimate represents some middle ground among the varying scientific options. Some think that the equivalent of the step from the emergence of trilobites to the domestication of fire goes like a shot in all planetary systems; others think that, even given ten or fifteen billion years, the evolution of a technical civilization is unlikely. This is not a subject on which we can do much experimentation as long as our investigations are limited to a single planet. Multiplying 9 UNCLASSI FIED/ /TOR OFFICIAL UO£ ONLY UNCLASSI FIED/ / FOR OFFICIAL USE ONLY these factors together, we find Ns fp ne fl fi fc~lx 109, a billion planets on which technical civilizations have arisen at least once. But that is very different from saying that there are a billion planets on which technical civilizations now exist. For this we must also estimate fL. W hat percentage of the lifetime of a planet is marked by a technical civilization? The Earth has harbored a technical civilization characterized by radio astronomy for only a few decades out of a lifetime of a few billion years. So far, then, for our planet fL is less than 1/108, a millionth of a percent. And it is hardly out of the question that we might destroy ourselves tomorrow. Suppose this were a typical case, and the destruction so complete that no other technical civilization - of the human or any other species - were able to emerge in the five or so billion years remaining before the Sun dies. Then Ns fp ne fl fi fc fL ~ 10, and, at a given time there would be only a tiny smattering, a handful, a pitiful few technical civilizations in the galaxy, the steady state number maintained as emerging societies replace those recently self-immolated. The number N might be even as small as 1 if civilizations tend to destroy themselves soon after reaching a technological phase; there might be no one for us to talk with but ourselves. And that we do but poorly. C ivilizations would take billions of years of tortuous evolution, and then snuff themselves out in an instant of unforgivable neglect. But consider the alternative, the prospect that at least some civilizations learn to live with technology; that the contradictions posed by the vagaries of past brain evolution are consciously resolved and do not lead to self destruction; or that, even if major disturbances occur, they are reveres in the subsequent billions of years of biological evolution. Such societies might live to a prosperous old age, their lifetimes measured perhaps on geological or stellar evolutionary time scales. If 1 percent of civilizations can survive technological adolescence, take the proper fork at this critical historical branch point and achieve maturity, then fL ~ 1/100, N ~ 107, and the number of extant civilizations in the galaxy is in the millions. Thus, for all our concern about the possible unreliability of our estimates of the early factors in the Drake equation, which involve astronomy, organic chemistry and evolutionary biology, the principal uncertainty comes to economics and politics and what, on Earth, we call human nature. It seems fairly clear that if self-destruction is not the overwhelmingly preponderant fate of galactic civilizations, then the sky is softly humming with messages from the stars. These estimates are stirring. They suggest that the receipt of a message from space is, even before we decode it, a profoundly hopeful sign. It means that someone has learned to live with high technology; that it is possible to survive technological adolescence. This alone, quite apart from the contents of the message, provides a powerful justification for the search for other civilizations. 4. The Drake Equation is Over-Simplified In the nearly fifty years (1961-2009) elapsed since Frank Drake proposed his equation, a number of scientists and writers tried to find out which numerical values of its seven independent variables are more realistic in agreement with our present-day knowledge. Thus there is a considerable amount of literature about the Drake equation nowadays, and, as one can easily imagine, the results obtained by the various authors largely differ from one another. In other words, the value of N, that various authors obtained by different assumptions about the astronomy, the biology and the sociology implied by the Drake equation, may range from a few tens (in the pessimist's view) to some 10 UNCLASSI FIED/ /FOR OFFICIAL USE ONEY UNCLASSIFIED/Zron OFFICIAL USE CM hY million or even billions in the optimist's opinion. A lot of uncertainty is thus affecting our knowledge of A/ as of 2010. In all cases, however, the final result about N has always been a sheer number, i.e., a positive integer number ranging from 1 to millions or billions. This is precisely the aspect of the Drake equation that this author regarded as "too simplistic" and improved mathematically in his paper #IAC -08-A4.1.4, entitled "The Statistical Drake Equation" and presented on O ctober 1st, 2008, at the 59th International Astronautical C ongress (IAC ) held in Glasgow, Scotland, UK, September 29th thru O ctober 3rd, 2008. That paper is attached herewith as Appendix B. Newcomers to SETI and to the Drake equation, however, may find that paper too difficult to be understood mathematically at a first reading. Thus, I shall now explain the content of that paper "by speaking easily." I thank the reader for his or her attention. 5. The Statistical Drake Equation W e start by an example. C onsider the first independent variable in the Drake equation (7), i.e., Ns, the number of stars in the Milky W ay galaxy. Astronomers tell us that approximately there should be about 350 millions stars in the galaxy. O f course, nobody has counted (or even seen in the photographic plates) all the stars in the galaxy! There are too many practical difficulties preventing us from doing so: just to name one, the dust clouds that don't allow us to see even the Galactic Bulge (i.e. the central region of the galaxy) in the visible light (although we may "see it" at radio frequencies like the famous neutral hydrogen line at 1420 MHz). So, it doesn't make any sense to say that Ns = 350 x 106, or, say (even worse) that the number of stars in the galaxy is (say) 354,233,321, or similar fanciful exact integer numbers. That is just silly and non-scientific. Much more scientific, on the contrary, is to say that the number of stars in the galaxy is 350 million plus or minus, say, 50 millions (or whatever values the astronomers may regard as more appropriate, since this is just an example to let the reader understand the difficulty). Thus, it makes sense to REPLAC E each of the seven independent variables in the Drake equation (7) by a MEAN V ALUE (350 millions, in the above example) PLUS O R MINUS A C ERTAIN STANDARD DEV IATIO N (50 millions, in the above example). By doing so, we have made a great step ahead: we have abandoned the too-simplistic equation (7) and replaced it by something more sophisticated and scientifically more serious: the STATISTIC AL Drake equation. In other words, we have transformed the classical and simplistic Drake equation (7) into an advanced statistical tool for the investigation of a host of facts hardly known to us in detail. In other words still: • W e replace each independent variable in (7) by a RANDO M V ARIABLE, labeled D, (from Drake). • W e assume that the MEAN V ALUE of each d , is the same numerical value previously attributed to the corresponding independent variable in (7). • But now we also ADD A STANDARD DEV IATIO N uDi on each side of the mean value, that is provided by the knowledge gathered by scientists in each discipline encompassed by each Di. 11 U N CLASSI FI E D/ / FOR OFFICIAL USE ONI Y U N CLASSI FI E D/ / FOR OFFICIAL USE ONLY Having so done, the next question is: How can we find out the PRO BABILITY DISTRIBUTIO N for each o,? For instance, shall that be a Gaussian, or what? This is a difficult question, for nobody knows, for instance, the probability distribution of the number of stars in the galaxy, not to mention the probability distribution of the other six variables in the Drake equation (7). There is a brilliant way to get around this difficulty, though. W e start by excluding the Gaussian because each variable in the Drake equation is a PO SITIV E (or, more precisely, a non-negative) random variable, while the Gaussian applies to REAL random variables only. So, the Gaussian is out. Then, one might consider the large class of well-studied and positive probability densities called "the gamma distributions," but it is then unclear why one should adopt the gamma distributions and not any other. The solution to this apparent conundrum comes from Shannon's Information Theory and a theorem that he proved in 1948: "The probability distribution having maximum entropy (= uncertainty) over any FINITE range of real values is the UNIFO RM distribution over that range," This is proven in Appendix A of the present document. So, at this point, we assume that each of the seven D: in (7) is a UNIFO RM random variable, whose mean value and standard deviation is known by the scientists working in the respective field (let it be astronomy, or biology, or sociology). Notice that, for such a uniform distribution, the knowledge of the mean value pD and of the standard deviation crD automatically determines the RANGE of that random variable in between its lower (called a,) and upper (called fy) limits: in fact these limits are given by the equations =Ad, ~V3crD. (the "surprising" factor 73 in the above equations comes from the definitions of mean value and standard deviation: please see equations (12), (15) and (17) in Appendix B for the relevant proof). So the uniform distribution of each random variable Di is perfectly determined by its mean value and standard deviation, and so are all its other properties. The next problem is the following: O K, since we now know everything about each uniformly distributed D,, what is the probability distribution of N , given that N is the product (7) of all then, ? In other words, not only do we want to find the analytical expression of the probability density function of N, but we also want to relate its mean value nN to all mean values //„ of the d, , and its standard deviation aN to all standard deviations 0) Mean value {N^e'2 V ariance 4=^/(/- Standard deviation — J- crlV = ^ e 2 ye'7 -1 All the moments, i.e. k-th moment k1 — ^k\ = e^ e 2 Mode (= abscissa of the lognormal peak) ^imde = flpeak “ ^ & V alue of the Mode Peak 1 —Zv(«n»dC)= f— -ef‘-e2 er Median (= fifty-fifty probability value for /V ) median = m = e^ Skewness K’, 4"' 4 (e°' - 1 f ^ff! + 3e2’2 + 6^ + bf Kurtosis 7%=^ +2?ff2+3e2aI -6 (^2) Expression of //in terms of the lower (a/) and upper {bi) limits of the Drake uniform input random variables Di =X^)=^^^4^d^^ j=I i=l ^i~ai Expression of o-2 in terms of the lower (a/) and upper (bi) limits of the Drake uniform input random variables Di = yff2 = y1 «A [info)-info)]2 M Y i M {^-a^ 7. An Example Explaining the Statistical Drake Equation To understand how things work in practice for the statistical Drake equation, please consider the following table 2. It is made up of three columns: • The first column on the left lists the seven input sheer numbers that also become • The mean values (middle column). • Finally the last column on the right lists the seven input standard deviations. 15 UNCLASSI FIED/ /FOR OFFICIAL USE OHLY UNCLASSIFIED//rOn OFFICIAL USE CM hY The bottom line is the classical Drake equation (7). W e see that, for this particular set of seven inputs, the classical Drake equation (i.e. the product of the seven numbers) yields a total of 3500 communicating extraterrestrial civilizations existing in the galaxy right now. Table 2. Input Values (i.e. mean values and standard deviations) for the Seven Drake Uniform Random Variables Di . The first column on the left lists the seven input sheer numbers that also become the mean values (middle column). Finally the last column on the right lists the seven input standard deviations. The bottom line is the classical Drake equation (7). The statistical Drake equation, however, provides a much more articulated answer than just the above sheer number /V = 3500. In fact, a MathC ad code written by this author and capable of performing all the numerical calculations required by the statistical Drake equation for a given set of seven input mean values plus seven input standard deviations, yields for N the lognormal distribution (thin curve) plotted in Figure 2. W e see immediately that the peak of this thin curve (i.e. the mode) falls at about «nBde = "peak =^6^' -250 (this is equation (99) of Appendix B), while the median (fifty­ fifty value splitting the lognormal density in two parts with equal undergoing areas) falls at about m,^^ = e1 V ^^-1 =11195 and so the expected number of/V may actually be even much higher than the 4590 provided by the mean value alone! The "upper limit of the one-sigma confidence interval" (as statisticians call it), i.e. the sum 4590+ 11195 = 15,785, yields a higher number still! (Note: the "lower limit of the one-sigma confidence interval is ZERO because the lognormal distribution is PO SITIV E (or, more correctly, non-negative)). Finally, the reader should note that the thick curve depicted in Figure 2 is just the NUMERIC AL solution of the statistical Drake equation for a FINITE number of 7 input factors. Figure 2 actually shows that this curve "is well interpolated" by the lognormal distribution (thin curve), i.e., by the neat analytical expression provided by the C entral Limit Theorem for an INFINITE number of factors in the Drake equation. That is, in conclusion, Figure 2 visually shows that taking 7 factors or an infinity of factors "is almost the same thing" already for a value as small as 7. 17 UNCLASSIFIED/^TOR OFFICIAL USE G Nh¥ UNCLASSIFIED/Zron OFFICIAL USE CM hY 8. Finding the Probability Distribution of the Et-Distance B y Virtue of the Statistical Drake Equation Having solved the statistical Drake equation by finding the lognormal distribution, we are now in a position to solve the ET-DISTANC E problem by resorting to statistics again, rather than just to the purely deterministic Distance Law (5), as we did in Section 2. This is "scientifically more serious" than just the purely deterministic Distance Law (5) inasmuch as the new statistical Distance Law will yield a PRO BABILITY DENSITY for the Distance, with the relevant mean value and standard deviation. In other words, the Distance Law (5) itself becomes a random variable whose probability distribution, mean value and standard deviation must be computed by "replacing" into (5) the fact that /V is now known to follow the lognormal distribution. This is mathematically described in detail in Section 7 of Appendix A. The important new result is the PRO BABILITY DENSITY FO R THE DISTANC E, the equation of which is holding for r>0. This is equation (114) of Appendix B. Starting from this equation, the MEAN V ALUE O F THE random variable ET_DISTANC E is computed as (ET_Distance) = C e 3 e18 (10) which is equation (119) of Appendix B, and finally the ET_DISTANC E STANDARD DEV IATIO N (ID which is equation (123) of Appendix B. O f course, all other descriptive statistical quantities, such as moments, cumulants etc. can be computed upon starting from the probability density (9), and the result is Table two hereafter, that is Table 3 of Appendix B. Finally, to complete this section, as well as this "introduction to the statistical Drake equation," the numerical values that equations (10) and (11) yield for the Input Table 1 are determined. They are, respectively: rm ean_vatue ^^ £ ~ 2,670 light yCSIS (12) UNCLASSI FIED/ /fOR OmCIAIi USE ONLY U N CLASSI FI E D/ / FOR OFFICIAL USE ONLY which is equation (153) of Appendix B, and ^ 47" J 47" ET-Disune =Ce 3 e18 V 9 -1» 1,309 light years (13) which is equation (154) of Appendix B. 19 UNCLASSI FIED//FOR OFFICIAL USE ONLY U N CLASSI FI E D/ / FOR OFFICIAL USE ONLY Table 2. Summary of the Properties of the Probability Distribution That Applies to the Random Variable ET_Distance Yielding the (average) Distance B etween Any Two Neighboring Communicating Civilizations in the G alaxy Random variable ET_Distance between any two neighboring ET civilizations in galaxy assuming they are UNIFO RMLY distributed throughout the whole galaxy volume. Probability distribution Unnamed Probability density function C 6 ^Qfl fa jy h(j(l [jy _ /ET Distanc V ) “ r- — £ r V 2.TC T 2 a2 Numerical constant C related to the Milky W ay size C = V 6 r g ^ ^laxy ~ ^^S light years Mean value (ET-Dislance} = C e 3 e18 V ariance 2 cr 2 ffETDta^^ 9 -1 Standard deviation ^ET.Distane = ^ 3 £ 18 Ie9 -1 All the moments, i.e. k-th moment (Er_Distance*) = C*/ ^e 18 Mode (= abscissa of the lognormal peak) ^mude = rpeak “ C £ € V alue of the Mode Peak Peak V alue of /ET_Distane(r) = “ ZET^Distanc C 'ncde ) — /”— ^ ^C V 2/T cr Median (= fifty-fifty probability value for N) median = m = C e 3 Skewness e-* 1 (kJ ( C 3 e 9 -4e CT2 5 <72 (T2 e 2 -3 e 18 + 2 e 6 5 g2 4 CT2 ). In fact, the probability density (9) has an infinite tail on the right, as clearly shown in Figure 3, and hence its mean value must be higher than its peak value. As given by (10) and (12), its mean value is rm ean m !ue = C e 3 e18 * 2670 light years. This is the MEAN (value of the) DISTANC E at which we can expect to find extraterrestrials. 22 UNCLASSI FIED/yrOR OFFICIAL USE ONLY U N CLASSI FI E D/ / FOR OFFICIAL USE ONLY After having found the above two distances (1933 and 2670 light years, respectively), the next natural question that arises is: "what is the range, back and forth around the mean value of the distance, within which we can expect to find extraterrestrials with "the highest hopes?" The answer to this question is given by the notion of standard deviation that we already found to be given by (11) and (13), ^ET-Disi™ =Ce 3 e18 ve 9 -1 -1309 light years. More precisely, this is the so-called 1-sigma (distance) level. Probability theory then shows that the nearest extraterrestrial civilization is expected to be located within this range, i.e. within the two distances of (2670-1309) = 1361 light years and (2670+ 1309) = 3979 light years, with probability given by the integral of /ET Distana.(r) taken in between these two lower and upper limits, that is: p3979light yean L, ' /mj^(r>*0.75 = 75% (15) 361 lightyears In plain words: with 75 percent probability, the nearest extraterrestrial civilization is located in between the distances of 1361 and 3979 light years from us, having assumed the input values to the Drake Equation given by table 1. If we change those input values, then ail the numbers change again, of course. 9. The "Data Enrichment Principle" as the B est CLT Consequence Upon the Statistical Drake Equation (Any Number of Factors Allowed) As a fitting climax to all the statistical equations developed so far, let us now state our "DATA ENRIC HMENT PRINC IPLE." It simply states that "The Higher the Number of Factors in the Statistical Drake equation, The Better." Put in this simple way, it simply looks like a new way of saying that the C LT lets the random variable Y approach the normal distribution when the number of terms in the sum (4) approaches infinity. And this is the case, indeed. 10. Conclusions W e have sought to extend the classical Drake equation to let it encompass Statistics and Probability. This approach appears to pave the way to future, more profound investigations intended not only to associate "error bars" to each factor in the Drake equation, but especially to increase the number of factors themselves. In fact, this seems to be the only way to incorporate into the Drake equation more and more new scientific information as soon as it becomes available. In the long run, the Statistical Drake equation might just become a huge computer code, growing in size and especially in the depth of the scientific information it contains. It would thus be Humanity's first "Encyclopaedia Galactica." 23 UNCLASSI FIED/ )Wn OFFICIAL U6E ONLY U N CLASSI FI E D/ / FOR OFFICIAL USE ONLY Unfortunately, to extend the Drake equation to Statistics, it was necessary to use a mathematical apparatus that is more sophisticated than just the simple product of seven numbers. 24 UNCLASSI FIED/ /FOR OFFICIAL M EE ONLY UNCLASSIFIED/Zron OFFICIAL USE CM hY Appendix A: Proof of Shannon's 1948 Theorem Stating That the Uniform Distribution is the "M ost Uncertain" One Over a Finite Range of Values Information Theory was initiated by C laude Shannon (1916-2001) in his well-known 1948 two papers: Rejmtri with carecncm non; 77* ^7 S^w Ttc'^icai Jbws#Z Vol. 27, pp r9-42Jt 623-656, July October 1943 A Mathematical Theory of Communication ByC. E SHANNON In this Appendix, we wish to draw attention to a couple of theorems that Shannon proves on pages 36 and 37 of his work, and read, respectively (note that Shannon omits the upper and lower limits of all integrals in the first theorem: they are minus infinity and plus infinity, respectively): 5. Let pl .r) be a coe-dimensional distribution The form ofp(x) sh rug a maximum entr opy' subject to the condition that the standar d deviation ofx be fixed at a is Gaussian. To show this we must maximize H(x) = -J p(x\lo^p{x)dx with ^’ = y p(x)x* dr and ^ = / P(x) dx as constraints* This requires, by the calculus of variations, maximizing ^ I -p(x) logp(x) + Ap(x)jr + pp(x)] dx. The condition for this is -1 - iogp(x) + Ax2 + p = 0 and consequently (adjusting the constants to satisfy the constraints) and 25 UNCLASSI FIED//FOR OFFICIAL USE ONLY UNCLASSIFIED//FOR OFFICIAL USE OM hY 7. If t is limited to a half line (p(x) = 0 for x < 0) and the hi st moment of x is fixed at a: a= I p(x)xdx. JO then the maximum entropy occurs when plx) = l^1/^ and is equal to log m. Now, we wish to point out that there is a third possible case, other than the two given by Shannon. This is the case when the probability density function p(*) >s limited to a FINITE INTERV AL an accthai ||CFnNly U N CLASSI FI E D/ / FOR OFFICIAL USE ONLY 31 U N CLASSI FI E D//FOR OFFICIAL USE OM I Y U N CLASSI FI E D/ / FOR OFFICIAL USE ONLY So, let us take the natural logs of both sides of the Statistical Drake equation (3) and change it into a sum: 7 ^D, f=l (4) It is now convenient to introduce eight new (positive) random variables defined as follows: J r = InQv) 7, =ln(£>,) j = 1,..,7. Upon inversion, the first equation of (5) yields the important equation, that will be used in the sequel N = eY . (6) We are now ready to take STEP THREE. STEP 3: TH E TRANSF ORM ATION LAW OF RAND OM VARIABLES So far we did not mention at all the problem: “which probability distribution shall we attach to each of the seven (positive) random variables D- ?” It is not easy to answer this question because we do not have the least scientific clue to what probability distributions fit at best to each of the seven points listed in Section 1. Yet, at least one trivial error must be avoided: claiming that each of those seven random variables must have a Gaussian (i.e. normal) distribution. In fact, the Gaussian distribution, having the well- known bell-shaped probability density function fx{x-^a)=-=^e 20) (7) V2^cr has its independent variable y ranging between -oo and oo and so it can apply to a real random variable F only, and never to positive random variables like those in the statistical Drake equation (3). Period. Searching again for probability density functions that represent positive random variables, an obvious choice would be the gamma distributions (see, for instance, ref. [6]). However, we discarded this choice too because of a different reason: please keep in mind that, according to (5), once we selected a particular type of probability density function (pdf) for the last seven of equations (5), then we must compute the (new and different) pdf of the logs of such random variables. And the pdf of these logs certainly is not gamma-type any more. It is high time now to remind the reader of a certain theorem that is proved in probability courses, but, unfortunately, does not seem to have a specific name. It is the transform ation law (so we shall call it, see, for instance, ref [5]) allowing us to compute the pdf of a certain new random variable F that is a known function y = ^(x) of another random variable X having a known pdf. In other words, if the pdf fx (x) of a certain random variable X is known, then the pdf /} (y) of the new random variable F, related to X by the functional relationship y=X*) <8) can be calculated according to this rule: 1) First invert the corresponding non-probabilistic equation y = g(x) and denote by xf(y) the various real roots resulting from the this inversion. 2) Second, take notice whether these real roots may be either finitely- or infinitely-many, according to the nature of the function v = ^(x). 3) Third, the probability density function of F is then given by the (finite or infinite) sum where the summation extends to all roots x( (y) and |^ (x/(>7)^ is the absolute value of the first derivative of g(x) where the /-th root x- (y) has been replaced instead of x. Since we must use this transformation law to transfer from the D, to the F- = ln(£)J, it is clear that we need to start from a D t pdf that is as simple as possible. The gamma pdf is not responding to this need because the analytic expression of the transformed pdf is very complicated (or, at least, it looked so to this author in the first instance). Also, the gamma distribution has two free parameters in it, and this “complicates” its application to the various meanings of the Drake equation. In conclusion, we discarded the gamma distributions and confined 32 UNCLASSI FIED/ /FOR OFFICIAL USE G Nh¥ U N CLASSI FI E D/ / FOR OFFICIAL USE ONLY ourselves to the simpler uniform distribution instead, as shown in the nest section. 4. STEP 4: ASSUM ING TH E EASIEST INPUT D ISTRIBUTION F OR EACH D i: TH E UNIF ORM D ISTRIBUTION L et us now suppose that each of the seven D i is distributed U NIF ORM L Y in the interval rang ing from the low er lim it ai >0 to the upper lim it b^ > ^. This is the same as saying that the probability density function of each of the seven Drake random variables D i has the equation ^ - U i) (a^ + a^i + b] ) aj + a^ + bf 3 (b; -a,) " 3 The second moment of the uniform distribution is thus uniform_Dj2 at + a^ + bi 3 (13) From (12 and (13) we may now derive the variance of the uniform distribution ^uniformDj = (uniform_Dj2) -(uniform_Dj)” /uniform D: CO ( ® i ~a i with 0<^ y — fly />Jin^(Z>J--2in(i»A+_2]-aJin^(a£^ ^ “ a i ...(32) The variance of Y i = ln(D/) is now given by (32) minus the square of (31), that, after a few reductions, yield: A 6)=4- *'=v,7 °i ~a i InfeA^lnfe) (28) 2 _ ’ _ 1 ^Mlnfe )-l"k F (33) P robability density functions of the natural log s of all the uniform ly distributed D rake random variables D i. This is indeed a positive function of y over the interval In («f) < y < In^ ), as for every pdf, and it is easy to see that its normalization condition is fulfilled: f^b ( \; H^ ^ j ^)_e'^) I,,, A w^ = I, / ,7— dy=—r----------=1 Jln(oj r Jln(af)£>f. -fly bi~a i ♦ 429) Next we want to find the mean value and standard deviation of K , since these play a crucial role for future developments. The mean value ^) is given by r^Jy*^ ,--------dy 'info) b{ - a^ ^M^kfMH (30) This is thus the m ean value of the natural log of all the uniform ly distributed D rake random variables D i {Y ^k^D ^^[ln(^)-l]-fly[ln(fly)-l] ^i _ a i - (31) Whence the corresponding standard deviation . _ «iMMMz>!iW. f w Let us now turn to another topic: the use of Fourier transforms, that, in probability theory, are called “characteristic functions,” Following again the notations of Papoulis (ref [5]) we call “characteristic function”, ^(f) , of an assigned probability distribution K , the Fourier transform of the relevant probability density function, that is (with j = -7-1 ) ^M’^V^AA. (35) J—x> The use of characteristic functions simplifies things greatly. For instance, the calculation of all moments of a known pdf becomes trivial if the relevant characteristic function is known, and greatly simplified also are the proofs of important theorems of statistics, like the Central Limit Theorem that we will use in Section 4. Another important result is that the characteristic function of the sum of a finite number of independent random variables is simply given by the product of the corresponding characteristic functions. This is just the case we are facing in the Statistical Drake equation (3) and so we are now led to find the characteristic function of the random variable T,, i.e. ^ ^=L/^A W^y = dy 36 UNCLASSI FIED/ ^QRQEEICTfil I ISF ONI Y U N CLASSI FI E D/ / FOR OFFICIAL USE ONLY = ^^(M (1+k )^ = ii bj - a I Jln(a() bi - Hi 1 + i£ ^J C M b,) _^l+XHa.) b'+ jC _a^K (bi -«;)(* + jc) (P i -«/)(! + K) (36) This author regrets that he was unable to compute the last integral analytically. He had to compute it num erically for the particular values of the 14 a; and bi that follow from Table 1 and equations 17. The result was the probability density function for F = ln(A9 plotted in the following Figure 2. Thus, the characteristic function of the natural log of the D rake uniform random variable D i is g iven by M<7= b^-a^ (bi-a^ + jt) (37) 3.3 STEP 7: F IND ING TH E PROBABILITY D ENSITY F UNCTION OF N, BUT ONLY NUM ERICALLY NOT ANALYTICALLY F igure 2. Probability density function of F = ln(W) computed numerically by virtue of the integral (39). The two “funny gaps” in the curve are due to the numeric limitations in the MathCad numeric solver that the author used for this numeric computation. Having found the characteristic functions Oy^) of the logs of the seven input random variables D i , we can now immediately find the characteristic function of the random variable F = ln(N) defined by (5). In fact, by virtue of (4), of the well-known Fourier transform property stating that “the Fourier transform of a convolution is the product of the Fourier transforms*’, and of (37), it immediately follows that Oy(^) equals the product of the seven 0r (c ): (bi-a^ + jZ ) ■ (38) We are now just one more step from finding the probability density of N, the number of ExtraTerrestrial Civilizations in the Galaxy predicted by our Statistical Drake equation (3). The point here is to transfer from the probability density function of F to that of N, knowing that F = ln(Nh or alternatively, that N=exp(Fh as stated by (6). We must thus resort to the transformation law of random variables (9) by setting y = ■?(*)=**• (40) The next step is to invert this Fourier transform in order to get the probability density function of the random variable F = ln(A9. In other words, we must compute the following inverse Fourier transform This, upon inversion, yields the sing le root ■*1(30=44=M?)' (41) On the other hand, differentiating (40) one gets fAyh^-fe^^A^ 1 «'W = «J and g (.v। (v)) = e1 n^ = y (42) where (41) was already used in the last step. The general transformation law (9) finally yields (bi-aA^+ jC ) dC . (39) ^,=Ww4 aW< (43 ) 37 UNCLASSI FIED//FOP fiFFTHfll IISF ONI Y UNCLASSIFIED//FOR OmCIAL USE OM LY This probability density function /^(y) was computed numerically by using (43) and the numeric curve given by (39), and the result is shown in Figure 3. N = Number of ET Civilizations in Galaxy F igure 3. The num eric (and not analytic) probability density function curve f^{y} of the number N of ExtraTerrestrial Civilizations in the Galaxy according to the Statistical Drake equation (3). We see that the curve peak (Le, the mode) is very close to low values of A, but the tail on the right is high, meaning that the resulting mean value (N^ is of the order of thousands. We now want to compute the mean value (A) of the probability density (43). Clearly, it is given by 00 (N) = jyfN(y)(ly. (44) 0 This integral too was computed numerically, and the result was a perfect m atch with A-3500 of (22), that is (JV) = 3499.99880 177509 + 0.00000012 4914686i (45) Note that this result was computed numerically in the complex domain because of the Fourier transforms, and that the real part is virtually 3500 (as expected) while the imaginary part is virtually zero because of the rounding errors. So, this result is excellent, and proves that the theory presented so far is mathematically correct. Finally we want to consider the standard deviation. This also had to be computed numerically, resulting in a N =3953.42910 143389 + 0.00000003 28000581 . (46) This standard deviation, higher than the mean value, implies that N might range in between 0 and 7453. This completes our study of the probability density function of A if the seven uniform Drake input random variable D , have the mean values and standard deviations listed in Table 1. We conclude that, unfortunately, even under the sim plifying assum ptions that the D i be uniform ly distributed^ it is im possible to solve the full problem analytically, since all calculations beyond equation (38) had to be perform ed num erically. This is no g ood. Shall we thus loose faith, and declare “impossible” the task of finding an analytic expression for the probability density function fN(y) ? Rather surprisingly, the answer is “no”, and there is indeed a way out of this dead-end, as we shall see in the next section. 5. TH E CENTRAL LIM IT TH EOREM (CLT) OF STATISTICS Indeed there is a good, approximating analytical expression for fN(y) , and this is the following log norm al probability density function fAy^^}=---^~e 2a' O^°) y y!2^a ■ (47) To understand why, we must resort to what is perhaps the most beautiful theorem of Statistics: the Central Limit Theorem (abbreviated CLT). Historically, the CLT was in fact proven first in 1901 by the Russian mathematician Alexandr Lyapunov (1857-1918), and later (1920) by the Finnish mathematician Jari Waldemar Lindeberg (1876-1932) under weaker conditions. These conditions are certainly fulfilled in the context of the Drake equation because of the “reality” of the astronomy, biology and sociology involved with it, and we are not going to discuss this point any further here. A good, synthetic description of the Central Limit Theorem (CLT) of Statistics is found at the Wikipedia site (ref. [7]) to which the reader is referred for more details, such as the equations for the Lyapunov and the Lindeberg conditions, making the theorem “rigorously” valid. 38 UNCLASSIFIED/>4=On OFFICIAL USE Ohl I Y UNCLASSIFIED//FOR OFFICIAL USE OM EY P ut in loose term s, the C L T states that, if one has a sum of random variables even NOT identically distributed, this sum tends to a norm al distribution w hen the num ber of term s m aking up the sum tends to infinity. Also, the norm al distribution m ean value is the sum of the m ean values of the addend random variables, and the norm al distribution variance is the sum of the variances of the addend random variables. Let us now write down the equations of the CLT in the form needed to apply it to our Statistical Drake equation (3). The idea is to apply the CLT to the sum of random variables given by (4) and (5) w hatever their probability distributions can possibly be. Tn other words, the CLT applied to the Statistical Drake equation (3) leads immediately to the following three equations: 1) The sum of the (arbitrarily distributed) independent random variables K makes up the new random variable K 2) The sum of their mean values makes up the new mean value of K 3) The sum of their variances makes up the new variance of K In equations: i=l 7 To understand this fact better in mathematical terms consider again of the transformation law (9) of random variables. The question is: what is the probability density function of the random variable N in equation (6), that is, what is the probability density function of the lognormal distribution? To find it, set y=g(4=^x- (49) This, upon inversion, yields the sing le root *i(y)=*(y)=in(y)- <5°) On the other hand, differentiating (49) one gets x’W^' and ^'(^(y^^^j (51) where (50) was already used in the last step. The general transformation law (9) finally yields ^ (4=Lfft^=n A ^bll • (32) < |sfoW)| M Therefore, replacing the probability density on the right by virtue of the well-known normal (or Gaussian) distribution given by equation (7), the lognormal distribution of equation (47) is found, and the derivation of the lognormal distribution from the normal distribution is proved. In view of future calculations, it is also useful to point out the so-called “Gaussian integral”, that is: This completes our synthetic description of the CLT for sum s of random variables. 6. TH E LOGNORM AL D ISTRIBTION IS TH E D ISTRIBUTION OF TH E NUM BER N OF EXTRATERRESTRIAL CIVILIZATIONS IN TH E GALAXY The C L T m ay of course be extended to products of random variables upon taking the log s of both sides, just as w e did in equation (3). It then follow s that the exponent random variable, like Y in (6), tends to a norm al random variable, and, as a consequence, it follow s that the base random variable, like N in (6), tends to a log norm al random variable. A>0, B = reaL (53) This follows immediately from the normalization condition of the Gaussian (7), that is J- e '^ dx=\, ^2k v (54) just upon expanding the square at the exponent and making the two replacements (we skip all steps) 39 UNCLASSIFIED//FOR OmCIAh USE ONLY U N CLASSI FI E D/ / FOR OFFICIAL USE ONLY In the sequel of this paper we shall denote the independent variable of the lognormal distribution (47) by a lower case letter n to remind the reader that corresponding random variable N is the positive integer number of ExtraTerrestrial Civilizations in the Galaxy, In other words, n will be treated as a positive real number in all calculations to follow because it is a “large” number (i.e. a continuous variable) compared to the only civilization that we know of, i.e. ourselves. In conclusion, from now on the log norm al probability density function of N w ill be w ritten as Upon setting £ = 0 into (56), the normalization condition for fN(n) follows f/w(»>= I. (59) Upon setting ^ = 1 into (56), the important m ean value of the random variable N is found (60) fN(n)=L .^^e ^ (»>0) « fixer (56) Having so said, we now turn to the statistical properties of the lognormal distribution (55), i.e. to the statistical properties that describe the number N of ExtraTerrestrial Civilizations in the Galaxy. Upon setting £ = 2 into (56), the mean value of the square of the random variable N is found ^N2^^ e2a'. (61) The variance of N now follows from the last two formulae: Our first goal is to prove an equation yielding all the moments of the lognormal distribution (56), that is, for every non-negative integer ^ = 0, 1,2, ... one has (62) (57) The relevant proof starts with the definition of the k- th moment One then transforms the above integral by virtue of the substitution The square root of this is the important standard deviation form ula for the N random variable The third moment is obtained upon setting £ = 3 into (56) (64) Finally, upon setting £ = 4, the fourth moment of N is found W4W^ . (65) ln[w]=z. (58) The new integral in z is then seen to reduce to the Gaussian integral (53) (we skip all steps here) and (57) follows Our next goal is to find the cumulants of N. In principle, we could compute all the cumulants K- from the generic r-th moment p, by virtue of the recursion formula (see ref [8]) Ki= A ” -1) Kk ^"~k ■ ^ 40 UNCLASSI FIED//FOR OFFICIAL UDE ONLY UNCLASSIFIED//FOR OFFICIAL USE OM LY In practice, however, here we shall confine ourselves to the computation of the first four cumulants only because they only are required to find the skewness and kurtosis of the distribution. Then, the first four cumulants in terms of the first four moments read: ^i = Ai ^2 = A2 _ ^ 1 ^=^-3^ K,-Kf K^=^-4 KX K3-3Kl-6K2 K?-Kf. These equations yield, respectively: CT2 Kx =e^. K, =?'/(/ -1) 9 2 A", = e3f,e2° . K4 = Z^2’’ (/-l)’ ^’! + 3e2ff! + 6 ^ +6, (67) (68) (69) (70) (71) From these we derive the skewness and the kurtosis J^=/-2 + 2^+3^ -6. fe) (73) Finally, we want to find the mode of the lognormal probability density function, i.e, the abscissa of its peak. To do so, we must first compute the derivative of the probability density function fN(n) of equation (56), and then set it equal to zero. This derivative is actually the derivative of the ratio of two functions of n, as it plainly appears from (57). Thus, let us set for a moment £-(„).M?Mi (74) 2a2 where “E” stands for “exponent,” Upon differentiating this, one gets E'(n) = -^-2(ln[n]-A)-. (75) 2cr « But the lognormal probability density function (56), by virtue of (74), now reads , \ 1 e~E^ ^W=-s—-— <76> V2^cr fi So that its derivative is ^/etj^^ 2j^2^LkfLlzl^ dr 4 14 <7 m2 = i -c^jH] . (77) 4 171(7 W2 Setting this derivative equal to zero means setting £'(/0^+ 1 = 0 (78) That is, upon replacing (75), -L(1hH-a)+1 = 0. (79) (7 Rearranging, this becomes Infra]—// + a2 =0 (80) and finally ^m&de ~ wpeak ^^ ^ (81) This is the m ost likely num ber of ExtraTerrestrial C ivilizations in the G alaxy. How likely? To find the value of the probability density function fN{n) corresponding to this value of the mode, we must obviously replace (81) into (56). After a few rearrangements, one then gets 41 UNCLASSIFIED/jtf OR OFFICIAL USE ONI Y U N CLASSI FI E D/ / FOR OFFICIAL USE ONLY 1 _ — /w(«nvde) = -^=-----e ^’^ 2 cr (82) (85) This is “how likely" the m ost likely num ber of ExtraTerrestrial C ivilizations in the G alaxy is, L e, it is the peak heig ht in the log norm al probability density function fN(n) * Then, after a few reductions that we skip for the sake of brevity, the full equation (83) is turned into Next to the mode, the median m (ref [9]) is one more statistical number used to characterize any probability distribution. It is defined as the independent variable abscissa m such that a realization of the random variable will take up a value lower than m with 50% probability or a value higher than m with 50% probability again. In other words, the median m splits up our probability density in exactly two equally probable parts. Since the probability of occurrence of the random event equals the area under its density curve (i.e. the definite integral under its density curve) then the median m (of the lognormal distribution, in this case) is defined as the integral upper limit m : fN{n}dn=rX— ^e ^ =f (82 2 2 (86) that is (87) Since from the definition (85) one obviously has erf(0)=0, (87) becomes whence finally (88) median = m = ep (89) In order to find m, we may not differentiate (83) with respect to tn, since the "precise” factor ’A on the right would then disappear into a zero. On the contrary, we may try to perform the obvious substitution 2 =^Mz^f 2a 2 z>0 (84) This is the m edian of the log norm al distribution of N. In other w ords, this is the num ber of ExtraTerrestrial civilizations in the G alaxy such that, w ith 50 % probability the actual value of N w ill be low er than this m edian, and w ith 50 % probability it w ill be hig her. In conclusion, we feel useful to summarize all the equations that we derived about the random variable N in the following Table 2. into the integral (83) to reduce it to the following integral defining the error function erf(z) Random variable N- number of communicating ET civilizations in Galaxy Probability distribution Lognormal Probability density function /v(„)=l._J_e 20) Mean value Variance a2=^Z^-l) Standard deviation 4) Now consider the sphere around us having a radius r. The volume of such as sphere is _ 4 f ET_Distance 'Uu r_Sphere ” ^ ^ i 2 (101) In the last equation, we had to divide the distance “ET Di stance” betw een ourselves and the nearest ET Civilization by 2 because we are now going to make the unwarranted assumption that all ET C ivilizations are equally space from each other in the G alaxy! This is a crazy assumption, clearly, and should be replaced by more scientifically- grounded assumptions as soon as we know more about our Galactic Neighbourhood. At the moment, however, this is the best guess that we can make, and so we shall take it for granted, although we are aware that this is weak point in the reasoning. Having thus assum ed that ET C ivilizations are U NIF ORM L Y SP AC ED IN THE G AL AXY , w e can w rite dow n this proportion: ^G alaxy ^Our _ Sphere N T That is, upon replacing both (100) and (101) into (102): 4 (ET-DistanceY ^G alaxy _ 3 y 2 / The only unknow n in the last equation is ET-Distance, and so w e m ay solve for it, thus g etting the: (AVERAG E) D ISTANC E B ETWEEN ANY P AIR OF NEIG HB OU RING C IVIL IZ ATIONS IN THE G AL AXY (104) where the positive constant C is defined by C^V^^^gZ^* 28845 light years.(105) Equations (104) and (105) are the starting point for our first application of the Statistical Drake equation, that we discuss in detail in the coming sections of this paper. PROBABILISTIC D ERIVATION OF TH E PROBABILITY D ENSITY F UNCTION F OR ET_D ISTANCE The probability density function (pdf) yielding the distance of the ET Civilization nearest to us in the Galaxy and presented in this section, was discovered by this author on September 5 th, 2007. He did not disclose it to other scientists until the SETI meeting run by the famous mathematical physicist and popular science author, Paul Davies, at the “Beyond” Center of the University of Arizona at Phoenix, on February 5-6-7-8, 2008. This meeting was also attended by SETI Institute experts Jill Tarter, Seth Shostak, Doug Vakoch, Tom Pierson and others. During this author’s talk, Paul Davies suggested to call “the Maccone distribution” the new probability density function that yields the ET_Distance and is derived in this section. 46 UNCLASSI FIED/ /FOR OmCIAh USE ONLY UNCLASSIFIED//FOR OFFICIAL USE OM EY Let us go back to equation (104). Since N is now a random variable (obeying the lognormal distribution), it follows that the ET_Distance must be a random variable as well. Hence it must have some unknown probability density function that we denote by /et_Distant(r) (106) where r is the new independent variable of such a probability distribution (it is denoted by r to remind the reader that it expresses the three­ dimensional radial distance separating us from the nearest ET civilization in a full spherical symmetry of the space around us). The question then is: what is the unknown probability distribution (106) of the ET_Distance? We can answer this question upon making the two formal substitutions Upon replacing (111) into (9), we then find This is the denominator of (9). The numerator simply is the lognormal probability density function (56) where the old independent variable x must now be re-written in terms of the new independent variable y by virtue of (109). By doing so, we finally arrive at the new probability density function fY (y) N>.x ET_distance -> y (107) into the transformation law (8) for random variables. As a consequence, (104) takes form Rearranging and replacing y by r, the final form is: y = ^(x) = -^ = C*x 3. vx (108) In order to find the unknown probability density /ET-DistanJ^ we nOW tO ^P^ the rule (9) t0 (108). First, notice that (108), when inverted to yield the various roots x^yX yields a sing le real root only ^v)^. (109) y Then, the summation in (9) reduces to one term only. Second, differentiating (108) one finds * (*)=“•* 3 ■ (110) Thus, the relevant absolute value reads Now, just replace C in (113) by virtue of (105). Then: We have discovered the probability density function yielding the probability of finding the nearest ExtraTerrestrial C ivilization in the G alaxy in the spherical shell betw een the distances r and r+ dr from Earth: (114) In t^C ukt.vy^aki.iy 2 JET Distant'^J — l~— er V2^cr holding for r > 0. STATISTICAL PROPERTIES OF TH IS D ISTRIBUTION 47 UNCLASSIFIED/>fOR OFFICIAL WEE ONLY UNCLASSIFIED//FOR OFFICIAL USE OM hY We now want to study this probability distribution in detail. Our next questions are: I) What is its mean value? 2) What are its variance and standard deviation? 3) What are its moments to any higher order? 4) What are its cumulants? 5) What are its skewness and kurtosis? 6) What are the coordinates of its peak, Le. the mode (peak abscissa) and its ordinate? 7) What is its median? The first three points in the list are all covered by the following theorem: all the moments of (113) are given by (here k is the generic and non­ negative integer exponent, i.e. k = 0,1,2,3,... > 0) Upon setting £ = 1 into (117), the important m ean value of the random variable ET_D istance is found (Er_Distance) = C e 3 e18 (11?) Upon setting £ = 2 into (117), the mean value of the square of the random variable ET_Distance is found ET_Dis rance2 ) = C2 c 3 ^9 (120) The variance of ET_Distance now follows from the last two formulae with a few reductions: °et Distanffi “ {ET_Dis rance } - (ET_Distance) ET_Distance*}= £ rk /ET_DiStmiE(r)^ 2 .V So, the variance of ET_D istance is (121) To prove this result, one first transforms the above integral by virtue of the substitution 2 £ _r2 9 °ET_Distan® “ ^ ^ ‘ e9 -1 (122) (116) The square root of this is the important standard deviation of the ETD istance random variable Then the new integral in z is then seen to reduce to the known Gaussian integral (53) and, after several reductions that we skip for the sake of brevity, (115) follows from (53). In other words, we have proven that (ET_DistanceA ) = CA e 3 e 18 (117) Upon setting A=0 into (117), the normalization condition for /ET Distant (f) follows ET_Distan®(r) Jr = 1. (118) ^^=<^“^^-1- (123) The third moment is obtained upon setting A - 3 into (117) ^ET-Distance3^ = C3 e~p e 2 . (124) Finally, upon setting ^ = 4 into (117), the fourth moment of ET_Distance is found 4 8^ J ET_Distance4^ = C4 ^ 3 t^9 . (125) 48 UNCLASSIFIED/jTOR OFFICIAL USE ONLY UNCLASSIFIED//FOR OFFICIAL USE OM hY Our next goal is to find the cumulants of the ET_Distance. In principle, we could compute all the cumulants A' from the generic Ath moment fii by virtue of the recursion formula (see ref. [8]) ^=^-ZL , k*^- <126>*=i ^ / e fl e 2 -3 e 18 +2 e 6 C3 e 9 -4e 9 -3e 9 +12 e 3 -6e 9 ...(132) and the kurtosis In practice, however, here we shall confine ourselves to the computation of the first four cumulants because they only are required to find the skewness and kurtosis of the distribution (113). Then, the first four cumulants in terms of the first four moments read: A i - Ai ^2 - A2—^r ^3=^-3^ K2-K? /4 = A4 -4K| Kj -3X22 -6^2 K^ - ^. (127) = C 9 +2e3 +3e 9 -6. (133) Next we want to find the mode of this distribution, i.e. the abscissa of its peak. To do so, we must first compute the derivative of the probability density function /ET Djstan2 on the right would then disappear into a zero. On the contrary, we may try to perform the obvious substitution This is the m ost likely ET_D istance/iwH Earth, How likely ? To find the value of the probability density function /ET Dist.m(C(r) corresponding to this value of the mode, we must obviously replace () into (). After a few rearrangements, which we skip for the sake of brevity, one gets (147) into the integral (146) to reduce it to the following integral (85) defining the error function erf(z)« Then, after a few reductions that we leave to the reader as an exercise, the full equation (145), defining the median, is turned into the corresponding equation involving the error function erflx) as defined by (85): Peak Value of ./ET_Distane(*) — /ET_Distance (4r»de ) 50 UNCLASSIFIED//^^ OFFICIAL U6E ONLY U N CLASSI FI E D/ / FOR OFFICIAL USE ONLY Table 3. Summary of the properties of the probability distribution that applies to the random variable ET_Distance yielding the (average) distance between any two neighboring communicating civilizations in the Galaxy. Random variable ET_Distance between any two neighboring ET Civilizations in Galaxy assuming they are UNIFORMLY distributed throughout the whole Galaxy volume. Probability distribution Unnamed (Paul Davies suggested “Maccone distribution”) Probability density function JET_Distai ^ln ' 2 V f3 J J (r) = r.__L—e 2°-2 n^-7 r~— cr V2^cr (Defining the positive numeric constant C) C = ^ ^lllt„ hG aia^ * 28845 light years Mean value (Er_Dis tance) = C e 3 e18 Variance 2 (72 ^ET.Distan® “ £ e9 -1 Standard deviation °ETJ>istan© “ 6 *e 1 AU the moments, i.e. ^th moment _^A k2 —ET.Distance^C* e 3e 18 Mode (= abscissa of the probability density function peak) ^nude = F peak — C € € Value of the Mode Peak Peak Value of /ET_DiSta111E(r) = . 3 - — ~ /eT_Distance Corrode ) ~ r~— £ £ CV2^ cr Median (= fifty-fifty probability value for ET_Distance) median = m = C e 3 Skewness (*J c3 a 2 Ser2 er2 e~^ e 2 -3e18 + 2 e 6 \ > So-2 5 a2 4 cr2 cy t.T_4^T^3(,T + 12eT - X 3 ZziV -6c 9 7 Kurtosis 4 er2 c2 2 <72 -£v = e“ +2cy+3c~-6 Expression of // in terms of the lower (a,) and upper ft) limits of the Drake uniform input random variables D i z^lM lM Expression of tr in terms of the lower (a/) and upper (b/) limits of the Drake uniform input random variables D , ^y^ y ^MM-h^)]2 it Y l it fe-^f 51 UNCLASSI FIED/ /*QR QEEICT M IISFOhllY UNCLASSIFIED//FOR OFFICIAL USE OM LY that is (148) (149) Since from the definition (147) one obviously has erf(O)=O, (149) yields (150) whence finally This is the m edian of the log norm al distribution of N. bi other w ords, this is the num ber of ExtraTerrestrial civilizations in the G alaxy such that, w ith 50 % probability the actual value of N w ill be low er than this m edian, and w ith 50 % probability it w ill be hig her. In conclusion, we feel useful to summarize all the equations that we derived about the random variable N in the following Table 2. NUM ERICAL EXAM PLE OF TH E ET_D ISTANCE D ISTRIBUTION In this section we provide a numerical example of the analytic calculations carried on so far. Consider the Drake Equation values reported in Table 1. Then, the graph of the corresponding probability density function of the nearest ET_Distance, /ET Distaiie(r), is shown in Figure 6* median -m = C e 3 (151) F igure 6. This is the probability of finding the nearest ExtraTerrestrial Civilization at the distance r from Earth (in light years) if the values assumed in the Drake Equation are those shown in Table L The relevant probability density function /ET.DistancW is given by equation (113). Its mode (peak abscissa) equals 1933 light years, but its mean value is higher since the curve has a high tail on the right: the mean value equals in 52 UNCLASSIFIED/^TOR OFFICIAL USE ONLY UNCLASSIFIED//FOR OFFICIAL USE OM hY fact 2670 light years. Finally, the standard deviation equals 1309 light years: THIS IS G OOD NEWS F OR SETL inasm uch as the nearest ET C ivilization m ig ht lie at just 1 sigma - 2670-1309 = 1361 light years from us. From Figure 6, we see that the probability of finding ExtraTerrestrials is practically zero up to a distance of about 500 light years from Earth. Then it starts increasing with the increasing distance from Earth, and reaches its maximum at rn^^rpeak=C e 3 e 9 *1933 light years (152) This is the M OST L IKEL Y VAL U E of the distance at w hich w e can expect to find the nearest ExtraTerrestrial civilization. It is not, however, the mean value of the probability distribution (113) for /ET oiStane(r). In fact, the probability density (113) has an infinite tail on the right, as clearly shown in Figure 6, and hence its mean value must be higher than its peak value. As given by (119), its mean value is rm ean_m iU e = Ce 3 e18 * 2670 light years. (153) This is the M EAN (value of the) D ISTANC E at w hich w e can expect to find Extra Terrestrials. After having found the above two distances (1933 and 2670 light years, respectively), the next natural question that arises is: “what is the range, forth and back around the mean value of the distance, within which we can expect to find ExtraTerrestrials with “the highest hopes ?,” The answer to this question is given by the notion of standard deviation, that we already found to be given by (123) ^ a 2 \ a 2 ^ETjiistane =Ce 3 ^18 V e 9 -1 * 1309 light years. ...(154) More precisely, this is the so called 1-sigma (distance) level. Probability theory then shows that the nearest ExtraTerrestrial civilization is expected to be located within this range, i.e. within the two distances of (2670-1309) = 1361 light years and (2670+1309) = 3979 light years, with probability given by the integral of /ET_DiStan