/*
 * readme.txt
 * ----------
 *
 * Description of program "gender.c".
 *
 * Copyright (c):
 * 2007:  Jrg MICHAEL, Adalbert-Stifter-Str. 11, 30655 Hannover, Germany
 *
 * SCCS: @(#) readme.txt  1.0  2007-07-11
 *
 * This file is subject to the GNU Lesser General Public License (LGPL)
 * (formerly known as GNU Library General Public Licence)
 * as published by the Free Software Foundation; either version 2 of the
 * License, or (at your option) any later version.
 * This file is distributed in the hope that it will be useful,
 * but WITHOUT ANY WARRANTY; without even the implied warranty of
 * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.
 *
 * You should have received a copy of the GNU Library General Public License
 * along with this file; if not, write to the
 * Free Software Foundation, Inc., 675 Mass Ave, Cambridge, MA 02139, USA.
 *
 * Actually, the LGPL is __less__ restrictive than the better known GNU General
 * Public License (GPL). See the GNU Library General Public License or the file
 * COPYING.LIB for more details and for a DISCLAIMER OF ALL WARRANTIES.
 *
 * There is one important restriction: If you modify this program in any way
 * (e.g. modify the underlying logic or translate this program into another
 * programming language), you must also release the changes under the terms
 * of the LGPL.
 * That means you have to give out the source code to your changes,
 * and a very good way to do so is mailing them to the address given below.
 * I think this is the best way to promote further development and use
 * of this software.
 *
 * If you have any remarks, feel free to e-mail to:
 *     ct@ct.heise.de
 *
 * The author's email address is:
 *    astro.joerg@googlemail.com
 */


This file contains some important information, which cannot be found
in the German "c't" article.

Index:

1. Overview of the program "gender"
2. The dictionary file "nam_dict.txt"
3. A few words on quality of data
4. What about Wikipedia / Wiktionary ?

5. The function "get_gender"
6. The function "check_nickname"
7. Other functions

8. Unicode chars
9. Statistics of first names in dictionary

10. Frequently Asked Questions
11. WWW resources for this program
12. History of the program


========================================================================


Overview of the program "gender"


The program "gender.c" is a program for determining the gender of a given
fist nane.

List of files:

a)  gen_ext.h     (contains macros and prototypes; may be changed)
b)  umlaut.h      (contains lists of umlauts)
c)  gender.c      (this is the "workhorse" of the program)
d)  nam_dict.txt  (dictionary file containing first names)

The file "nam_dict.txt", which contains a list of first names, uses the
char set "iso8859-1".

If you want to use "gender.c" as a library, delete the line
"#define GENDER_EXECUTABLE"  from the file "gen_ext.h".


========================================================================


The dictionary file "nam_dict.txt"


The program "gender.c" uses the dictionary file "nam_dict.txt" as a data
source. This file contains a list of more than 40,000 first names and
gender, plus some 600 pairs of "equivalent" names.
This list should be able to cover the vast majority of first names
in all European contries and in some overseas countries (e.g. China,
India, Japan, U.S.A.) as well.

Also included in this file is information on the approximate frequency
of each name. The scale goes from 1 (=rare) to 13 (=extremely common).
The value 10 has been formatted to represent at least 2 percent of
the population. (The values 11 to 13 have been added last.)
The scale is logarithmic. For countries with very good statistics,
each step (down to frequency 2) represents a factor of 2.
For example, a frequency value of 7 means that the correspondig first
name has an absolute in the range of 0.25 % to 0.5 %.

The sorting order of the file "nam_dict.txt" is governed by the search
algorithm of the program "gender.c". Hence, names with "expandable"
umlauts can be found twice in this dictionary, first with sorting
according to "expanded" umlauts, and second with sorting according to
"compressed" umlauts (e.g. '' is sorted like "Oe" and 'O').

You don't have to reformat this file for use in a unix environment,
because the DOS linefeeds (trailing '\r') are ignored when the file
is read.


========================================================================


A few words on quality of data


The dictionary of first names has been prepared with utmost care.

For example, the Turkish, Indian and Korean names in this dictionary
have all been independently lassified by several native speakers.
I also took special care to list only those names which can currently
be found.

The lesson from this?
Any modifications should be done very cautiously (and they must also
adhere to the sorting required by the search algorithm).

For example, knowing that "Sascha" is a boy's name in Germany, the author
never assumed the English "Sasha" to be a girl's name.
Knowing that "Jan" is a boy's name in Germany, I never assumed it to be
also a English short form of "Janet". Another case in point is the name
"Esra". This is a boy's name in Germany, but a girl's name in Turkey.

Or consider the following first names:

Ildik     female Hungarian name
Mitja      male Russian name
Elizaveta  rare name; looks like misspelled "Elizabeta"
Roelf      rare name; looks like German "Rolf" with an erroneous 'e'

Borchert, Oltmann, Sievert, Hartmann    look like common German surnames


========================================================================


What about Wikipedia / Wiktionary ?

The dictionary of first names (file "nam_dict.txt") is subject to the
GNU Free Documentation License. Therefore, anybody can use this file to
e.g. create national lists of first names in Wiktionary.

As you can see in a comment on
  http://en.wikipedia.org/wiki/List_of_Italian_given_names
it might be more appropriate to add any lists of first names to Wiktionary
rather than to Wikipedia.

Call  "gender -print_names_of_country  <country>  -unicode_file=<file>"
for this purpose. The suboption "-unicode_file=<file>" generates a native
unicode file, so you don't have to reformat any internal "replacement-chars"
(array "umlauts_unicode[]") for non-iso8859-1 umlauts.

Of course, you also have to give proper credit to author and data source.
The data source should at least be identified by:
    File "nam_dict.txt" from www.heise.de/ct, soft-link 0717182
(To avoid typos, it is best to copy and paste this line.)

Some additional remarks:

a)
It is quite possible that several people will want to work on names of
the same country.

Therefore, if you add the names of a given country to Wiktionary,
you should (before and afterwards) check for conflicting pices of work
(i.e. other people doing the same).

The probably easiest way might be to do a Google search for (e.g.)
"nam_dict.txt" and "0717182".
You see, that's why I urge you to identify the data source by:
    File "nam_dict.txt" from www.heise.de/ct, soft-link 0717182

b)
The following list might need special cleanup, because I found even some
of the most important names to be missing:
  http://it.wiktionary.org/wiki/Utente:Ft1bot/Antroponimi

c)
There is not a unique well-defined rule for transcriptions from cyrillic
alphabets. Many cyrillic names have more than one valid transcription,
and the transcription rules also vary from country to country.

For example, the following names refer to the same "original" cyrillic
name:
 - Aleksander, Aleksandr, Alexander, Alexandr, (and sometimes even Oleksandr)
 - Maria, Marija and Mariya
 - Tatiana, Tatjana and Tatyana
 - Vladimir and Wladimir

Since this dictionary emphasizes the true frequency of a name,
I strived to give all valid "normal" transcriptions equal or at least
similar frequency values.

Hence, if you prepare names lists for Russia, Belarus, Ukraine or
Bulgaria, you have to familarize yourself with the subleties of the
various cyrillic alphabets.

d)
Last but not least, Chinese and Korean names need special formatting.
For both countries, I have used a plus char ('+') "inside" a name to
symbolize '-', ' ' or an empty string. Thus, "Jun+Wei" represents the
names "Jun-Wei", "Jun Wei" and "Junwei".

In practice however, most Chinese first names appear as a "single" name,
while Korean first names are mostly written as separate names.
Therefore, the Chinese name "Xiao+Wei" should be formatted as "Xiaowei",
while the Korean name "Kyung+Hee" should be formatted as "Kyung Hee".


========================================================================


The function "get_gender"


The function "get_gender" is used to check first names and determine their
gender. This function has the following return values, which are defined
in the file "gen_ext.h":

IS_FEMALE         :  female first name
IS_MOSTLY_FEMALE  :  mostly female first name
IS_MALE           :  male first name
IS_MOSTLY_MALE    :  mostly male first name
UNISEX_NAME       :  unisex name (can be any gender)

NAME_NOT_FOUND    :  name not found
ERROR_IN_NAME     :  name contains an error
ERROR_IN_DATAFILE :  error in dictionary file "nam_dict.txt"

In the case of a "multiple" first name (e.g. "Carl J.D."), "get_gender"
checks every single name and combines the results. For example, the name
"A. Paul" is marked as male, because "A." can be any gender and "Paul"
is a boy's name.

An exception are some names like (e.g.) "Jean-Claude". This name is
directly labeled as "IS_MALE", because there is a corresponding entry
in the dictionary file.


========================================================================


The function "check_nickname"


The function "check_nickname" is used to check whether two first names
are "equivalent" (e.g. "Bill" and "William").
This function has the following return values, which are defined in the
file "gen_ext.h":

EQUIVALENT_NAMES  :  names are equivalent
NOT_EQUAL_NAMES   :  names are not equivalent

ERROR_IN_NAME     :  name contains an error
ERROR_IN_DATAFILE :  error in dictionary file "nam_dict.txt"


========================================================================


Other functions


Cleanup function

Upon starting the program, the dictionary file "nam_dict.txt" is statically
opened to speed up data access.

The down side is that you have to close this file when you finally exit
the program. Call the function "cleanup_gender()" for this task.



Consistency check

Call the program "gender" with argument "-check_consistency" to do a
consistency check. There should be no inconsistencies, of course.



Print all names of a given country

Call:   gender -print_names_of_country  <country>  [ -unicode_file=<file>" ]
The optional suboption "-unicode_file=<file>" generates a native unicode
file, so you don't have to reformat any internal "replacement-chars"
(array "umlauts_unicode[]") for non-iso8859-1 umlauts.



Print statistics

Call the program "gender" with argument "-statistics" or "-statistics -full"
to get elementary or full statistics, respectively.
Printing full statistics may take a while...


========================================================================


Unicode chars


This programme uses the char set "iso8859-1".
"Non-iso" unicode chars are represented as follows:
   256 = <A/>
   257 = <a/>
   258 = <>
   259 = <>
   260 = <A,>
   261 = <a,>
   262 = <C>
   263 = <c>
   268 = <C^> or <CH>
   269 = <c^> or <ch>
   271 = <d>
   272 = <> or <DJ>
   273 = <> or <dj>
   274 = <E/>
   275 = <e/>
   278 = <E>
   279 = <e>
   280 = <E,>
   281 = <e,>
   282 = <>
   283 = <>
   286 = <G^>
   287 = <g^>
   290 = <G,>
   291 = <g>
   298 = <I/>
   299 = <i/>
   304 = <I>
   305 = <i>
   310 = <K,>
   311 = <k,>
   315 = <L,>
   316 = <l,>
   317 = <L>
   318 = <l>
   321 = <L/>
   322 = <l/>
   325 = <N,>
   326 = <n,>
   327 = <N^>
   328 = <n^>
   336 = <>
   337 = <>
   338 =  or <OE>
   339 =  or <oe>
   344 = <R^>
   345 = <r^>
   350 = <S,>
   351 = <s,>
   352 =  or <S^> or <SCH> or <SH>
   353 =  or <s^> or <sch> or <sh>
   354 = <T,>
   355 = <t,>
   357 = <t>
   362 = <U/>
   363 = <u/>
   366 = <U>
   367 = <u>
   370 = <U,>
   371 = <u,>
   379 = <Z>
   380 = <z>
   381 = <Z^>
   382 = <z^>

A plus char ('+') "inside" a name symbolizes a '-', ' ' or an empty string
(this option applies to Chinese and Korean names only).
Thus, "Jun+Wei" represents the names "Jun-Wei", "Jun Wei" and "Junwei".


========================================================================


Statistics of first names in dictionary


Number of first names   :   41287
   girl's names         :   15400
   boy's  names         :   16100
   unisex names         :    9787
Equivalent name pairs   :     657

Names from Great Britain  :  2229
Names from Ireland        :   877
Names from U.S.A.         :  3902
Names from France         :  1817
Names from Italy          :  2830
Names from Malta          :   526
Names from Portugal       :  1040
Names from Spain          :  1639
Names from Belgium        :  1498
Names from Luxembourg     :   530
Names from the Netherlands:  3263
Names from East Frisia    :  2343   ("Ostfriesland")
Names from Germany        :  1953
Names from Austria        :  1722
Names from Swiss          :  2430
Names from Iceland        :  1150
Names from Denmark        :  1143
Names from Norway         :  1063
Names from Sweden         :   931
Names from Finland        :   842
Names from Estonia        :  1120
Names from Latvia         :   707
Names from Lithuania      :   845
Names from Poland         :   359
Names from Czech Republic :   405
Names from Slovakia       :   428
Names from Hungary        :   371
Names from Romania        :  1959
Names from Bulgaria       :  1790
Names from Bosnia and Herzegovina:  1210
Names from Croatia        :   897
Names from Kosovo         :   692
Names from Macedonia      :   838
Names from Montenegro     :   621
Names from Serbia         :   728
Names from Slovenia       :   635
Names from Albania        :  1694
Names from Greece         :   773
Names from Russia         :   480
Names from Belarus        :   438
Names from Moldova        :   403
Names from Ukraine        :   445
Names from ex-U.S.S.R. (Asian):   950
Names from Turkey         :  1358
Names from Arabia/Persia  :  1523
Names from Israel         :   660
Names from China          :  7334
Names from India/Sri Lanka:  1491
Names from Japan          :  1380
Names from Korea          :  1376
Names from Vietnam        :   307


Statistics for Great Britain  (statistics are very good):
                rare            medium          common         total
first  names:  1465  218 143 100 100  80  80  30 10  3  0 0 0   2229
girl's names:   619  103  92  51  57  46  43  15  3  1  0 0 0   1030
boy's  names:   642   87  35  37  33  26  34  14  7  2  0 0 0    917
unisex names:   204   28  16  12  10   8   3   1  0  0  0 0 0    282

Statistics for Ireland  (statistics are good):
                rare            medium          common         total
first  names:   312  134 124  78  80  65  53  25  6  0  0 0 0    877
girl's names:   176   72  60  43  46  44  27  16  2  0  0 0 0    486
boy's  names:   112   54  49  29  32  19  24   9  4  0  0 0 0    332
unisex names:    24    8  15   6   2   2   2   0  0  0  0 0 0     59

Statistics for U.S.A.  (statistics are very good):
                rare            medium          common         total
first  names:  2612  421 292 202 190 103  56  18  8  0  0 0 0   3902
girl's names:  1455  214 155  93  92  48  27   8  1  0  0 0 0   2093
boy's  names:   939  160  99  71  69  37  22  10  7  0  0 0 0   1414
unisex names:   218   47  38  38  29  18   7   0  0  0  0 0 0    395

Statistics for France  (statistics are very good):
                rare            medium          common         total
first  names:  1140  162 113 102  94  89  75  36  4  0  2 0 0   1817
girl's names:   690  107  72  57  60  58  43  13  0  0  1 0 0   1101
boy's  names:   314   41  35  41  33  27  32  22  3  0  1 0 0    549
unisex names:   136   14   6   4   1   4   0   1  1  0  0 0 0    167

Statistics for Italy  (statistics are very good):
                rare            medium          common         total
first  names:  1846  311 206 162 134  94  42  26  5  3  1 0 0   2830
girl's names:   884  155 118 101  75  45  21  12  1  1  1 0 0   1414
boy's  names:   957  155  88  61  59  49  21  14  4  2  0 0 0   1410
unisex names:     5    1   0   0   0   0   0   0  0  0  0 0 0      6

Statistics for Malta  (statistics are medium quality):
                rare            medium          common         total
first  names:   136   94  74  79  47  43  30  11  7  3  2 0 0    526
girl's names:    40   42  17  35  23  23  18   2  2  1  1 0 0    204
boy's  names:    89   49  50  40  20  20  12   9  5  2  1 0 0    297
unisex names:     7    3   7   4   4   0   0   0  0  0  0 0 0     25

Statistics for Portugal  (statistics are good):
                rare            medium          common         total
first  names:   508  181  84 120  70  36  14  13  7  3  3 1 0   1040
girl's names:   272   87  38  60  40  20   8   3  2  0  0 1 0    531
boy's  names:   235   92  46  58  28  16   6  10  5  3  3 0 0    502
unisex names:     1    2   0   2   2   0   0   0  0  0  0 0 0      7

Statistics for Spain  (statistics are very good):
                rare            medium          common         total
first  names:   942  214 180 106  78  56  30  16  9  6  1 1 0   1639
girl's names:   522  115  91  46  45  34  10   9  3  1  0 1 0    877
boy's  names:   411   98  88  60  30  21  19   6  6  5  1 0 0    745
unisex names:     9    1   1   0   3   1   1   1  0  0  0 0 0     17

Statistics for Belgium  (statistics are very good):
                rare            medium          common         total
first  names:   520  206 211 182 181 117  53  24  1  3  0 0 0   1498
girl's names:   291  127 127  92 101  58  26   7  0  2  0 0 0    831
boy's  names:   179   64  71  81  72  51  24  16  1  1  0 0 0    560
unisex names:    50   15  13   9   8   8   3   1  0  0  0 0 0    107

Statistics for Luxembourg  (statistics are very good):
                rare            medium          common         total
first  names:    84   67  74  69  61  74  54  33 11  1  2 0 0    530
girl's names:    41   26  38  40  44  48  33  13  2  0  1 0 0    286
boy's  names:    31   38  35  25  16  23  20  18  8  1  1 0 0    216
unisex names:    12    3   1   4   1   3   1   2  1  0  0 0 0     28

Statistics for the Netherlands  (statistics are very good):
                rare            medium          common         total
first  names:  1887  476 330 233 146  99  67  21  3  1  0 0 0   3263
girl's names:   955  249 186 120  74  58  28   9  0  0  0 0 0   1679
boy's  names:   756  168 106  91  52  28  35  12  3  1  0 0 0   1252
unisex names:   176   59  38  22  20  13   4   0  0  0  0 0 0    332

Statistics for East Frisia  (statistics are good):
                rare            medium          common         total
first  names:  1370  379 138 142 100  96  79  36  3  0  0 0 0   2343
girl's names:   619  122  54  68  62  57  45  15  0  0  0 0 0   1042
boy's  names:   604  210  75  69  34  35  32  19  3  0  0 0 0   1081
unisex names:   147   47   9   5   4   4   2   2  0  0  0 0 0    220

Statistics for Germany  (statistics are very good):
                rare            medium          common         total
first  names:  1193  185 159 123  88  79  75  44  7  0  0 0 0   1953
girl's names:   511  104  92  69  48  38  47  26  0  0  0 0 0    935
boy's  names:   661   73  63  54  39  40  27  18  7  0  0 0 0    982
unisex names:    21    8   4   0   1   1   1   0  0  0  0 0 0     36

Statistics for Austria  (statistics are very good):
                rare            medium          common         total
first  names:  1046  186 126 100  85  69  59  40  7  4  0 0 0   1722
girl's names:   630  120  78  59  51  47  37  16  3  1  0 0 0   1042
boy's  names:   397   63  43  40  34  22  22  24  4  3  0 0 0    652
unisex names:    19    3   5   1   0   0   0   0  0  0  0 0 0     28

Statistics for Swiss  (statistics are very good):
                rare            medium          common         total
first  names:  1230  347 289 204 171  99  53  32  5  0  0 0 0   2430
girl's names:   705  200 161 114 101  52  27  12  2  0  0 0 0   1374
boy's  names:   498  135 125  87  67  45  24  20  3  0  0 0 0   1004
unisex names:    27   12   3   3   3   2   2   0  0  0  0 0 0     52

Statistics for Iceland  (statistics are very good):
                rare            medium          common         total
first  names:   335  194 176 123 118 103  64  23 11  3  0 0 0   1150
girl's names:   180   97  86  53  59  43  38   7  7  1  0 0 0    571
boy's  names:   155   97  90  70  59  60  26  16  4  2  0 0 0    579
unisex names:     0    0   0   0   0   0   0   0  0  0  0 0 0      0

Statistics for Denmark  (statistics are very good):
                rare            medium          common         total
first  names:   320  204 175 127  95  85  58  55 23  1  0 0 0   1143
girl's names:   184  116  97  71  47  54  31  31  9  0  0 0 0    640
boy's  names:   131   84  73  53  48  30  27  24 14  1  0 0 0    485
unisex names:     5    4   5   3   0   1   0   0  0  0  0 0 0     18

Statistics for Norway  (statistics are good):
                rare            medium          common         total
first  names:   412  183 142  89  89  69  59  13  7  0  0 0 0   1063
girl's names:   234   92  73  46  45  32  31   3  1  0  0 0 0    557
boy's  names:   169   88  69  43  44  36  28  10  6  0  0 0 0    493
unisex names:     9    3   0   0   0   1   0   0  0  0  0 0 0     13

Statistics for Sweden  (statistics are very good):
                rare            medium          common         total
first  names:   290  139 117 118  96  62  44  44 20  1  0 0 0    931
girl's names:   177   82  55  67  53  31  23  19 10  0  0 0 0    517
boy's  names:   107   54  60  49  42  31  21  25 10  1  0 0 0    400
unisex names:     6    3   2   2   1   0   0   0  0  0  0 0 0     14

Statistics for Finland  (statistics are good):
                rare            medium          common         total
first  names:   215   99 129 113 107  72  60  33 13  1  0 0 0    842
girl's names:   126   57  64  60  58  30  35  13  8  0  0 0 0    451
boy's  names:    87   41  64  52  48  42  25  20  5  1  0 0 0    385
unisex names:     2    1   1   1   1   0   0   0  0  0  0 0 0      6

Statistics for Estonia  (statistics are good):
                rare            medium          common         total
first  names:   449  183 103 121  93  93  56  22  0  0  0 0 0   1120
girl's names:   232  105  49  66  47  47  29  11  0  0  0 0 0    586
boy's  names:   214   78  54  55  45  46  27  11  0  0  0 0 0    530
unisex names:     3    0   0   0   1   0   0   0  0  0  0 0 0      4

Statistics for Latvia  (statistics are good):
                rare            medium          common         total
first  names:   118  119  76 115 108  83  52  28  6  1  1 0 0    707
girl's names:    76   63  30  60  65  45  26  18  2  0  0 0 0    385
boy's  names:    42   56  46  55  43  38  26  10  4  1  1 0 0    322
unisex names:     0    0   0   0   0   0   0   0  0  0  0 0 0      0

Statistics for Lithuania  (statistics are good):
                rare            medium          common         total
first  names:   399  103  89  52  70  57  43  18  8  6  0 0 0    845
girl's names:   132   49  41  21  31  29  21   8  5  2  0 0 0    339
boy's  names:   267   54  48  31  39  28  22  10  3  4  0 0 0    506
unisex names:     0    0   0   0   0   0   0   0  0  0  0 0 0      0

Statistics for Poland  (statistics are good):
                rare            medium          common         total
first  names:    94   40  57  27  39  26  25  24 19  7  1 0 0    359
girl's names:    50   22  32  14  22  11   6  13  9  5  1 0 0    185
boy's  names:    44   18  25  13  17  15  19  11 10  2  0 0 0    174
unisex names:     0    0   0   0   0   0   0   0  0  0  0 0 0      0

Statistics for Czech Republic  (statistics are good):
                rare            medium          common         total
first  names:   124   55  51  27  45  34  30  17  8 13  1 0 0    405
girl's names:    76   30  31  13  27  19  19  13  4  2  0 0 0    234
boy's  names:    47   25  20  14  18  15  11   4  4 11  1 0 0    170
unisex names:     1    0   0   0   0   0   0   0  0  0  0 0 0      1

Statistics for Slovakia  (statistics are medium quality):
                rare            medium          common         total
first  names:   127   64  43  44  27  52  34  22  8  6  1 0 0    428
girl's names:    67   36  17  30  19  27  21   9  3  2  1 0 0    232
boy's  names:    60   28  26  14   8  25  13  13  5  4  0 0 0    196
unisex names:     0    0   0   0   0   0   0   0  0  0  0 0 0      0

Statistics for Hungary  (statistics are good):
                rare            medium          common         total
first  names:   117   49  20  34  37  36  29  21 21  7  0 0 0    371
girl's names:    65   28  11  17  18  26  20  15  9  2  0 0 0    211
boy's  names:    52   21   8  17  19  10   9   6 12  5  0 0 0    159
unisex names:     0    0   1   0   0   0   0   0  0  0  0 0 0      1

Statistics for Romania  (statistics are very good):
                rare            medium          common         total
first  names:  1084  415 157  94  57  59  49  27 10  5  2 0 0   1959
girl's names:   720  241  99  55  36  38  25  14  4  0  1 0 0   1233
boy's  names:   364  174  58  39  21  21  24  13  6  5  1 0 0    726
unisex names:     0    0   0   0   0   0   0   0  0  0  0 0 0      0

Statistics for Bulgaria  (statistics are good):
                rare            medium          common         total
first  names:   745  241 244 142 162 126  66  42 16  6  0 0 0   1790
girl's names:   432  144 137  81  87  73  49  24  6  3  0 0 0   1036
boy's  names:   311   97 105  61  75  53  17  18 10  3  0 0 0    750
unisex names:     2    0   2   0   0   0   0   0  0  0  0 0 0      4

Statistics for Bosnia and Herzegovina  (statistics are medium quality):
                rare            medium          common         total
first  names:   466  195 124 144 120  95  55  10  1  0  0 0 0   1210
girl's names:   205   85  59  45  51  37  28   7  1  0  0 0 0    518
boy's  names:   258  110  64  94  66  57  27   3  0  0  0 0 0    679
unisex names:     3    0   1   5   3   1   0   0  0  0  0 0 0     13

Statistics for Croatia  (statistics are good):
                rare            medium          common         total
first  names:   259  217 103  96  63  68  51  31  6  2  1 0 0    897
girl's names:   130   95  44  48  36  35  26  17  1  1  0 0 0    433
boy's  names:   126  122  58  44  26  33  25  14  5  1  1 0 0    455
unisex names:     3    0   1   4   1   0   0   0  0  0  0 0 0      9

Statistics for Kosovo  (statistics are medium quality):
                rare            medium          common         total
first  names:   271  121  67  90  63  54  23   3  0  0  0 0 0    692
girl's names:   145   43  25  37  28  23  14   3  0  0  0 0 0    318
boy's  names:   124   77  42  51  32  30   9   0  0  0  0 0 0    365
unisex names:     2    1   0   2   3   1   0   0  0  0  0 0 0      9

Statistics for Macedonia  (statistics are medium quality):
                rare            medium          common         total
first  names:   347  169  68  73  47  71  40  16  7  0  0 0 0    838
girl's names:   143   70  29  18  14  32  17  11  4  0  0 0 0    338
boy's  names:   201   97  39  54  32  39  23   5  3  0  0 0 0    493
unisex names:     3    2   0   1   1   0   0   0  0  0  0 0 0      7

Statistics for Montenegro  (statistics are medium quality):
                rare            medium          common         total
first  names:   133  119  68  85  59  77  44  25 10  1  0 0 0    621
girl's names:    58   51  30  31  20  31  20  12  6  0  0 0 0    259
boy's  names:    73   67  37  53  39  46  23  13  4  1  0 0 0    356
unisex names:     2    1   1   1   0   0   1   0  0  0  0 0 0      6

Statistics for Serbia  (statistics are medium quality):
                rare            medium          common         total
first  names:   290  103  45  55  82  68  47  26 10  2  0 0 0    728
girl's names:   150   49  19  21  31  32  22  13  5  1  0 0 0    343
boy's  names:   136   54  25  33  50  36  24  13  5  1  0 0 0    377
unisex names:     4    0   1   1   1   0   1   0  0  0  0 0 0      8

Statistics for Slovenia  (statistics are medium quality):
                rare            medium          common         total
first  names:   205  106  64  53  58  62  48  31  7  1  0 0 0    635
girl's names:    99   44  31  19  19  36  25  13  2  1  0 0 0    289
boy's  names:   105   62  32  33  36  26  22  18  5  0  0 0 0    339
unisex names:     1    0   1   1   3   0   1   0  0  0  0 0 0      7

Statistics for Albania  (statistics are good):
                rare            medium          common         total
first  names:   554  223 303 248 181 104  61  19  1  0  0 0 0   1694
girl's names:   293   91  86  74  60  49  34  12  1  0  0 0 0    700
boy's  names:   257  128 216 173 121  55  27   7  0  0  0 0 0    984
unisex names:     4    4   1   1   0   0   0   0  0  0  0 0 0     10

Statistics for Greece  (statistics are good):
                rare            medium          common         total
first  names:   354  123  49 107  45  32  27  18 11  5  2 0 0    773
girl's names:   240   74  30  58  25  20  16   9  7  1  1 0 0    481
boy's  names:   114   49  19  49  20  12  11   9  4  4  1 0 0    292
unisex names:     0    0   0   0   0   0   0   0  0  0  0 0 0      0

Statistics for Russia  (statistics are good):
                rare            medium          common         total
first  names:   115   69  48  40  31  55  34  33 28 23  4 0 0    480
girl's names:    70   38  27  26  19  29  14  14 11  8  2 0 0    258
boy's  names:    38   31  21  14  12  26  20  19 17 15  2 0 0    215
unisex names:     7    0   0   0   0   0   0   0  0  0  0 0 0      7

Statistics for Belarus  (statistics are medium quality):
                rare            medium          common         total
first  names:   133   64  49  35  27  19  29  37 22 14  9 0 0    438
girl's names:    80   31  22  17  17   7  12  20  6  5  5 0 0    222
boy's  names:    51   33  27  16  10  12  17  17 16  9  4 0 0    212
unisex names:     2    0   0   2   0   0   0   0  0  0  0 0 0      4

Statistics for Moldova  (statistics are medium quality):
                rare            medium          common         total
first  names:   147   73  24  32  29  23  31  21 14  9  0 0 0    403
girl's names:    90   39  12  12  18   7  19   8  6  4  0 0 0    215
boy's  names:    57   34  12  20  11  16  12  13  8  5  0 0 0    188
unisex names:     0    0   0   0   0   0   0   0  0  0  0 0 0      0

Statistics for Ukraine  (statistics are medium quality):
                rare            medium          common         total
first  names:   120   59  36  45  33  32  40  32 23 20  5 0 0    445
girl's names:    69   31  24  31  18  17  21   9 11 11  2 0 0    244
boy's  names:    50   25  12  14  15  15  19  23 12  9  3 0 0    197
unisex names:     1    3   0   0   0   0   0   0  0  0  0 0 0      4

Statistics for ex-U.S.S.R. (Asian)  (statistics are medium quality):
                rare            medium          common         total
first  names:   425  178 110 104  49  40  31   8  5  0  0 0 0    950
girl's names:   198   89  43  44  18  20  14   7  4  0  0 0 0    437
boy's  names:   223   88  66  59  31  20  17   1  1  0  0 0 0    506
unisex names:     4    1   1   1   0   0   0   0  0  0  0 0 0      7

Statistics for Turkey  (statistics are good):
                rare            medium          common         total
first  names:   220  406 209 224 155  93  37   6  7  1  0 0 0   1358
girl's names:    93  182  89  99  74  55  18   2  3  0  0 0 0    615
boy's  names:   109  195 105 113  74  33  17   4  4  1  0 0 0    655
unisex names:    18   29  15  12   7   5   2   0  0  0  0 0 0     88

Statistics for Arabia/Persia  (statistics are good):
                rare            medium          common         total
first  names:   396  488 254 168 102  68  37   6  4  0  0 0 0   1523
girl's names:   200  164  87  54  49  38  16   2  0  0  0 0 0    610
boy's  names:   188  316 164 113  51  29  20   4  4  0  0 0 0    889
unisex names:     8    8   3   1   2   1   1   0  0  0  0 0 0     24

Statistics for Israel  (statistics are medium quality):
                rare            medium          common         total
first  names:   156  152  75  73  49  85  59   8  3  0  0 0 0    660
girl's names:    55   67  32  23  13  40  33   4  2  0  0 0 0    269
boy's  names:    94   81  40  48  35  43  24   4  1  0  0 0 0    370
unisex names:     7    4   3   2   1   2   2   0  0  0  0 0 0     21

Statistics for China  (statistics are very good):
                rare            medium          common         total
first  names:  5319 1126 507 194  82  46  26  23  8  3  0 0 0   7334
girl's names:     0    0   0   0   0   0   0   0  0  0  0 0 0      0
boy's  names:     0    0   0   0   0   0   0   0  0  0  0 0 0      0
unisex names:  5319 1126 507 194  82  46  26  23  8  3  0 0 0   7334

Statistics for India/Sri Lanka  (statistics are good):
                rare            medium          common         total
first  names:   271  455 369 219 116  45  15   1  0  0  0 0 0   1491
girl's names:   154  131 120  76  47  19   6   0  0  0  0 0 0    553
boy's  names:    93  288 197 102  39  11   4   1  0  0  0 0 0    735
unisex names:    24   36  52  41  30  15   5   0  0  0  0 0 0    203

Statistics for Japan  (statistics are good):
                rare            medium          common         total
first  names:   426  265 204 142 153 109  57  22  2  0  0 0 0   1380
girl's names:   180   70  64  55  47  47  29  13  1  0  0 0 0    506
boy's  names:   213  178 130  77  98  55  22   5  1  0  0 0 0    779
unisex names:    33   17  10  10   8   7   6   4  0  0  0 0 0     95

Statistics for Korea  (statistics are good):
                rare            medium          common         total
first  names:    65  579 341 136  91  50  36  33 20 22  2 1 0   1376
girl's names:     6  200 115  43  25   8   0   0  0  1  0 0 0    398
boy's  names:    24  274 125  40   6   0   0   0  0  0  0 0 0    469
unisex names:    35  105 101  53  60  42  36  33 20 21  2 1 0    509

Statistics for Vietnam  (statistics are good):
                rare            medium          common         total
first  names:    32    0  36  41  39  39  34  31 26 19  7 1 2    307
girl's names:     0    0   0   0   0   0   0   0  3  0  0 0 1      4
boy's  names:     0    0   0   0   0   0   0   1  1  1  0 0 1      4
unisex names:    32    0  36  41  39  39  34  30 22 18  7 1 0    299



========================================================================


Frequently Asked Questions


1.
Is it possible to use this program in a commercial software project?

This is exactly what the Library GPL has been designed for.
See the file "COPYING.LIB" for details.

Probably the "safest" way to comply with the rules of the LGPL is
to put the program in a separate library (e.g. Windows-DLL).
In this way, only this library is subject to the LGPL and the "rest"
of the project still has the "old" rights of its owner, so you don't
have to give away your source code.


2.
Is a Java version of this program available?

The company who has written a Java version of my program "phonet"
has also promised to develop a native Java version of "gender" under
the LGPL - see:  https://opensource.softmethod.de/trac/opensource
The project name will be "gender4j".

Alternatively, you can also write a wrapper class in Java which uses
the Java Native Interface to call a C library.
There is an excellent article (in German) telling you how to do it:
"Kaffee mit Vitamin C", c't, issue 20/2000, pp.242-247
or:  www.heise.de/ct, soft-link 0020242.


3.
What is the speed of this program?

Due to the coherent use of binary search algorithms, this program runs
very fast even on an old computer.

In order to get measurable running times, you have to do tens or even
hundreds of thousands of function calls.


========================================================================


WWW resources for this program

www.heise.de/ct, soft-link 0717182



Unicode chars:

Can be looked up in Wikipedia.


========================================================================


History of the program

2007-05-23:  The first version of this program is submitted for
             publication in a German computer magazine (c't).
2007-08-06:  Version 1.0 is published in c't.


========================================================================
(End of file "readme.txt")
