Showing posts with label linguistics. Show all posts
Showing posts with label linguistics. Show all posts

Sunday, 31 January 2016

Confuci-us

Prerequisite: none

As a Chinese, the translation of "孔夫子" (kongˇ-foo-zhi˙) to "Confucius" (kon-fyoo-shus) has always been somewhat confusing.. very kon-fyoo-sing indeed! The last syllable has puzzled many a Chinese. Why would "zhi˙" become "shus"?

I thought this quirky translation was unique to Confucius until I found out yesterday that 孟子 (mengˋ-zhi˙) is also known as "Mencius" (men-shus). It finally clicked when I searched "Suncius" (孫子, author of The Art of War) and it showed up in "Vicipaedia".


The "-us" at the end is a masculine suffix. Julius. Marcus. Brutus. Confuci-us. As for pronounciation in Ecclesiastical Latin:

1) C before O is a hard C (as in k)
2) U is pronounced "oo"
3) C before I is a soft C (as in ch)

The last U is probably a short "oo". So instead of "kon-fyoo-shus", it should be:
"kon-foo-chi-us"

And leads me to wonder.. what is my Latinized Archaic name?

First of all, Archaic Chinese. Females more commonly used "氏" (shiˋ), which they attached to their father's or husband's surname. Problem is, my last name is from my mother's side (special case) and I am not married yet. "子" (zhi˙) was used for respected or scholarly people, although the fact that it was more explicitly used for males was due to the lack of female scholars.

There are more choices but they get very specific about status and age, and are not convenient for public use. For my case I better use "子".

My surname is 孫, so 孫子.
I wrote The Art of War, whoopee~

On the Latin part, pick any female Latin suffix of choice. Let me see..

Suncia
Suncilla
Suncilia
Suncilea
Suncina
Suncilina
Suncianna
Suncissa
Suncietta
Sunciella

I would not go for more than three syllables because that just sounds too princess-like.

Suncia
Suncilla
Suncina
Suncissa

Maybe something not so "flowery".

Suncia
Suncissa

Ehh.. I would rather not name myself after some furniture. "Cissa" is actually a genus of magpies, but preferable.

Suncissa

Tuesday, 29 December 2015

Nederlands

Prerequisite: none

Hallo!

The Dutch language is classified under Indo European, Germanic, West Germanic, Low Saxon Low Franconian, Low Franconian. English parts at the West Germanic branch, falling into the English category.

This post is only meant to provide a taste of the Dutch language. The grammar is very general and many exceptions are not noted. If you would like to learn conversational Dutch, Duolingo is a great free site to do so.

Here are some Dutch pronunciations given in IPA:


To aid whole words, Google Translate works pretty well.

The Familiar

Nominative pronoun:


*Sometimes je is an unemphasized jij, ze for zij, and we for wij. U is formal.

Verb conjugation:



Possessive pronoun:


Question (vraag):


Demonstrative pronoun:


Number (nummer):



Family (familie):


Colour (kleur):

red (rode)
orange (oranje)
yellow (gele)
green (groene)
blue (blauwe)
purple (paarse)
pink (roze)
brown (bruin)
white (witte)
grey (grijse)
black (zwarte)

Example phrases:

What is that?
(What is that?)
Wat is dat?

Thank you too
(Thank you too)
Dank je ook

I will be hungry and thirsty
(I shall have hunger and thirst)
Ik zul heb honger en dorst

My hand hurts
(My hand does hurt)
Mijn hand doet pijn

The pink pig likes to eat red apples
(The pink pig likes red apples to eat)
Het roze varken houdt van rode appels te eten

What would you like?
(What may it be?)
Wat mag het zijn?

Do we need spoons or forks?
(Have we spoons or forks need?)
Hebben wij lepels of vorken nodig?

They come from The Netherlands
(They come out Netherlands)
Zij kommen uit Nederland

The Unfamiliar

De / Het:

De and het are the "the" articles used for gender and neuter nouns. Here is a general guide to when to use which. I like to think of het as "it". The best way is not to memorize, but to accept the Dutch culture of what is considered "the" or "it", gendered or neutral.


The sandwiches (de boterhammen)
The man and woman (de man en de vrouw)
The artist (de kunstenaar)

The writing (het schrijven)
The cup (het kopje)
The book (het boek)

Lig / Zit / Sta:

Use these verbs when describing where something is. Here is a general guide to which verb to use.


The papers lie between the boxes.
De papieren liggen tussen de dozen.
My jacket lies under the bed.
Mijn jas ligt onder de bed.
A (dead) dog lies on the street.
Een hond ligt op de straat.

There sit women in the house.
Er zitten vrouwen in het huis.
The cat sits on the table.
De kat zit op de tafel.
Yuck, raisins sit in my bread.
Yuck, rozijnen zitten in mijn brood.

The lamp stands nearby the bookshelf.
De lamp staat nabij de boekenplank.
The buildings stand near the city.
De gebrouwen staan dichtbij de stadt.
There stands food in the kitchen.
Er staan eten in de keuken.

Interesting words of note:

zwembad = swim bath (swimming pool)
tijdschrift = time writing (magazine)
ziekenhuis = sick house (hospital)
dierentuin = animal garden (zoo)
schildpad = shield toad (turtle)
neushoorn = nose horn (rhinoceros)
vliegveld = fly field (airport)
hoofdstad = head city (capital)

This is only a sparse quarter of the Duolingo course. Still trying to get comfortable with some very Dutch words and patterns such as om te and er. Might update if I ever do.

Doei~

Saturday, 31 October 2015

Learning Languages

Prerequisite: none

My incentive for learning languages is not for conversation (me conversing with a stranger? What a joke!) but for the linguistics. For the sake of language itself. The way it is and its relation to other languages. As of now I know Chinese and English very well, and minimal Japanese, Dutch, Italian, Irish, Turkish, and maybe some Thai.

And also for curiosity. What is masculine or feminine noun? A gender or neuter noun? What are cases? When to use which cases? What is an agglutinative language? An eclipsis? A lenition? Go find out!

Duolingo

Great site to proceed at your own pace. What I like most about it is the convenience of definitions and audio. The interface design is very simple and there are many languages available. And why not, it is completely free. Completely free of additional purchases.

Readlang

The web browser feature similar to Google Translate but waaaaay better. You can read a page in foreign language and click on words you do not know for its definition. No need to flip the dictionary a billion times.

Invest in a language now~

Wednesday, 21 October 2015

Language Decryption

Followup of Code Decryption

Prerequisite: none

The Code Book by Simon Singh has a nice chapter on cracking ancient languages.


People first assumed that hieroglyphs were semantic pictographs and ideographs, and nothing more. No one bothered to challenge the assumption since the Ancient Egyptians were supposedly too "primitive" to come up with a phonetic system.

Then Rosetta Stone came around which contained hieroglyphs, demotic, and Greek on one slab, making a convenient crib, except that the Ancient Egyptian language has not been spoken for centuries. When Thomas Young spotted a cartouche on the Rosetta Stone, he suspected that it signified a pharaoh's name and that hieroglyphs might actually be phonetic.


He considered the historical context of several artifacts and associated names with cartouches. He could then deduce the sound values of each character. But his idea died down when he convinced himself that the alphabet was only applied to foreign names. Even with a collection of sounds, it did not seem to make meaning in regular text. At least, it made no sense to him..


Jean-François Champollion came across a cartouche. He figured that the the repeated letters are probably the repeated "s" in "Ramses". Being fluent in Coptic, he further suspected that the circumpunct reads "ra" as a rebus image.


And it worked. Ramses. After much substitution, it turns out the Egyptian hieroglyphs represented an ancestor of the Coptic language where some characters are phonetic and some are semantic. When Champollion traveled to Egypt he could really read hieroglyphs. Read-read hieroglyphs. Read. Hieroglyphs.


The Linear B tablet was found on Crete so the first speculation was that it is in Greek. But many Greek words end in "s", and the lack of a common last letter refutes that. Since the consensus was that the tablet contained a lost Minoan language, there was not much deciphering effort.


Alice Kober noted that there are around 100 characters, too much to be alphabetic and too few to be logographic, which makes it syllabic. She also noticed commonly occurring root words and suffixes, indicating an inflective language. It allowed her to associate syllables with the same consonants. Take Japanese as an example (except that Linear B had longer root words):

かく --> かきます
kaku      kakimasu
よむ --> よみます
yomu     yomimasu
つくる --> つくります
tsukuru      tsukurimasu
あぶ --> あびます
abu        abimasu
ぬぐ --> ぬぎます
nugu      nugimasu

Kober did the same analysis and grouped the Linear B characters by consonant, although she did not know what the consonants were.

Michael Ventris examined Kober's work and considered the geographical context of the tablet. He associated a regularly appearing word with "Knossos" and used it as a crib to identify other words such as "Pylos". Soon, he had enough cribs to substitute most of the text and fill in the gaps himself. The text was indeed in Greek, although there were some words he could not recognize. John Chadwick further identified the language as a kind of Archaic Greek. The ending "s" was dropped as a convention.

Monday, 12 October 2015

Code Decryption

Prerequisite: algebra

Been reading The Code Book: The Science of Secrecy from Ancient Egypt to Quantum Cryptography by Simon Singh. It is a very fascinating read between history, cryptography, and linguistics. In this post I compile some deciphering techniques. Some are simple while others are pure genius. But first, some terminology.

plaintext: original message (notated in lowercase)
ciphertext: enciphered message (notated in capitals)
algorithm: the method of enciphering a message
key: the premise of an enciphering method


A person needs to know both the algorithm and the key in order to decipher a ciphertext. But in many cases, the algorithm is obvious and the key can be traced from it.

Alphabetic Substitution

This is the simplest of algorithms where the alphabet is scrambled up to make a cipheralphabet. A good knowledge of English (or whatever the plaintext is written in) is enough to crack the ciphertext.

If a lone alphabet appears commonly throughout a ciphertext, one can deduce that is either "a" or "i". Similarly, a recurrence of a three letter cipherword is probably "and" or "the". If there is no vowel in a four letter cluster, one of the letters is probably a "y". The letter after the "q" must be a "u". If the spaces are eliminated, one can still guess common suffixes for a start. Lingual rules provide many handholds to decipherment.

Know your spelling rules, substitute what you can, and play a little hangman until you get the whole plaintext. That was how I cracked the Gnommish Alphabet in the Artemis Fowl series back in seventh grade.

Caesar Shift

Actually I lied. The Caesar Shift is even simpler. It shifts the alphabet several places, then uses it as the cipheralphabet. The number of shifts is agreed with the recipient beforehand.


This encipherment was used for extremely short messages, such as one phrase. There are not enough clues to reason with, but this is still a weak cipher considering that one only needs to test twenty six cipheralphabets at most to reach the plaintext. If that sounds like a lot of work to you, read on. You will much rather confront a Caesar Shift.

Vigenère Cipher

"The Indecipherable Cipher" utilizes the Vigenère Square, which is essentially all possible Caesar Shifts lined systematically to make a square:


What happens is that the sender and receiver agree on a keyword, such as "BLUE". To encipher a message, the first letter would be enciphered with the Caesar Shift starting with "B", the second letter with the shift starting with "L", the third with the shift starting with "U", the fourth with "E", and the fifth with "B" again. So the message "pig is hungry" enciphered with the keyword "BLUE" will be "QTAMTSORHCS".

B L U E B L U E B L U
p  i  g  i  s  h u n g  r  y
Q T A MT S O RH C S

This enciphering technique is a polyalphabetic cipher, which alternates between more than one cipheralphabet. This makes it harder to pick out letters by frequency as opposed to a monoalphabetic cipher, where you can almost guess correctly that the most common cipherletter probably represents the plainletter "e".

Charles Babbage figured that frequency analysis plays a big role concerning the nature of Caesar Shifts in the Vigenère Cipher. Arabs first came up with frequency analysis, the association of cipherletters with plainletters by occurrence. What happens is you get an graph describing the frequency distribution of alphabets in a language..


..then compare it to the frequency distribution of cipherletters in your ciphertext (similarly can be done with Zipf's Law for whole words). Match corresponding frequencies of letters and cipherletters, substitute, do some tweaking, and you should have the plaintext. This technique is not particularly significant for general Alphabetic Substitution since logical reasoning is enough to crack the cipher, but it gives a handhold in Vigenère decipherment.

Homophonic Cipher

The previous ciphers were especially vulnerable to letter frequencies. To make up for that, the homophonic cipher uses numbers as the cipheralphabet, and adds more cipherletters to even out the frequencies. Each cipherletter should appear just as often as another.


It takes much more thought to crack this cipher, but it is still possible. One can consider spelling rules, estimate the amount of extra cipherletters for each plainletter, and take both into account. There is much more trial and error, but it is not such a horror compared to the next cipher..

The Enigma

This is where encryption escalates quickly. Why read my words when you can see for yourself? This video on the Enigma Machine tells what you need to know.

Hooooo, what is this monster? From left to right, the machine components are lamp letters, keyboard, plugboard, first scrambler, second scrambler, third scrambler, and reflector. If you trace this diagram carefully, hitting the "C" key gives the output "F".


The Germans with their Enigma Machines changed their agreed scrambler setting everyday in order to securely encipher the scrambler setting of their actual messages. So a person receiving a message would set their Enigma Machine to the agreed day setting, decipher the new scrambler setting, set to the new setting, and then proceed to decipher the actual message. The plugboard setting stays the same.

The Machine is the algorithm and the scrambler setting is the key. The key is six letters long, where the first three are the starting letters of the scramblers and the last three is a repetition. The key in plaintext "pigpig" may be enciphered as "GXWLDN".

To obtain the key, Marian Rejewski cleverly mapped out the chains of letter relations. He analyzed numerous keys of one day setting and paired up the first and fourth letters of each six letter key, since they are repetitions of each other. A relation for A, B, C, D, E, and F can be:

A B C D E F
D A F E B C

Then he organized this relation into chains. In this example there are two separate chains:

four links: A --> D --> E --> B --> A
two links: C --> F --> C

The significance of this organization is that the plugboard cannot interfere with the amount of chains or links, so that it is useful for cracking the scrambler setting. Instead of finding one key among ten thousand million million keys, he had only 105,456 possible chain-link characteristics to consider. And then there was the manual labour of recording the number of chains and links for each scrambler setting, but they did it. It took a year.

When the Germans found out about their flaw they stopped repeating their keys. To find the key, Alan Turing used Rejewski's idea on cribs. A crib is a ciphertext in which you know its plaintext as well, which in their case had to be guessed. A common crib that the Germans provided was "weather" (or "wetter" in German) in their weather reports.

Then he had to figure some plugboard settings as well. The trial and error went something like this:

He had some data based on a crib.
Given that ciphertext "A" is plaintext "b",
assume that "A" is connected to "S" on plugboard.

A --> plugboard --> S --> scramblers --> ? --> plugboard --> b
A --> plugboard --> S --> scramblers --> F --> plugboard --> b

So he joined "A" to "S" on the plugboard, then saw which letter lighted up. If "b" lighted up, it shows that "b" is not connected to any letter on the plugboard. If "F" lighted up, which he knew should be "b", he then assumed that "F" is connected to "b".

All that is good, but what happens when there is a contradiction? Say, three letters seem to be plugged to each other. Yes, those would be incorrect deductions from an incorrect assumption. To turn mistakes into an advantage, Turing realized that all these incorrect deductions are definitely incorrect, and do not need to be tested further. I am still trying to get my head around this one.

And he threw these testings into a bombe.


For more historical context alongside the logic, you ought to read The Code Book. It is pleasant for leisure as well as for study. If this article makes you insecure about your internet privacy, I have another post coming up that will calm your nerves. Wait for it~

Friday, 14 August 2015

Omniglot

Prerequisite: none

Wondrous site: http://omniglot.com/

"The online encyclopaedia of writing systems and languages".

So what?

I knew that many minor languages exist, but this collection blew my mind. There are sooooo many languages and scripts within the neighbourhood of one region. It is also interesting to compare between proximally close languages.

Above all, foreign scripts are just so pretty. How many of us are aware that there are writing such as these in northern China? Or even that Manchurians have their own script?


Some pages to get started:

http://omniglot.com/writing/types.htm
Difference between abjad, abugida, alphabet, syllabary, and semanto-phoenetic writing systems.

http://omniglot.com/language/articles/index.htm
Articles, articles, and articles. About anything really.

http://omniglot.com/writing/direction.htm
Writing directions can be so odd you just have to see it to believe it.

The Ancient Berber script (in Morocco) runs from bottom to top.. hurhur!


Language is one of the many aspects of life that we take for granted. It is strange to think of the voices we never get to hear, of the records of past civilizations that are no longer intelligible, or simply legacies that no longer exist.. as well as current cultures that are heading down this path.