Friday, 16 October 2009

Google's tower of Babel

Is machine translation all that it’s cracked up to be?
Translators have been toiling away with computer-aided translation systems for over a decade now. The idea behind it is really quite straightforward. Text translated by a translator is stored in a memory bank, not as a complete document, but as chunks of text, that is, individual words, clauses, sentences or paragraphs. If a (near) identical chunk of text appears in a subsequent document for translation, the CAT software will suggest the saved chunk of text as the most ideal option for the translator to use. Of course, in order to work effectively, you need a memory bank that has been built up over a prolonged period of time, containing thousands and thousands of these chunks. Key to its practicability is the accrual of text. And while we may wish to call it machine translation technology, in fact to be effective in the first place, all the donkey work has to be done by humans. After all, at the end of the day, it’s the translator who settles on the most fitting translation. CAT has its limitations: it only really comes into its own when extremely repetitive texts have to be translated. Its main purpose is to ensure consistency, rather than to take any short-cuts.
Google has not been slow to cotton on to this concept and you can see why. By having made web-based data available in other languages, information is now at the fingertips of a much greater number of users worldwide. To do this, researchers at Google have been trawling the web and matching thousands, if not millions of documents and web-pages which have already been translated professionally (by humans). All this data has been entered into a databank to generate huge volumes of multilingual texts, this in addition to more primitive machine-translation methods (which translate single words rather than phrases). All this data is then used as a source for their translation tool. It would be hard for me as a professional translator to knock their motives in doing this; after all, making information available to a wider audience worldwide is a laudable goal, even though their motives may be first and foremost commercial. I only have to think of the frustration I’ve suffered trying to make heads or tails of all those French, Portuguese and Spanish sites I’ve browsed over the years. The results language-wise may be pretty woeful, but Google deserve a bit of a thumbs-up for their efforts in rolling back the frontiers.

There are lots of buts here though. If I’m trying to convert a legal contract from one language to another by machine translation, the likelihood is, if it’s been written in legalese, it will come out as pretty scrambled in the target language and thus be rendered fairly useless to anyone wishing to read and interpret it seriously. Only a fool would sign on the dotted line and make it legally binding. And as a customer, I would expect any documentation or product specifications I receive from a foreign supplier to be nothing less than word-perfect in my own language. Even everyday vernacular, used for example in films or documentaries, would fail the test: the inflexibility of machine tools in translating idioms and colloquialisms used for subtitles would result in some fairly bizarre interpretations. We can reflect on many more situations in real-life where machine translation just wouldn’t work: a job application, a poem, a product manual, a love letter, a menu, instructions for medication, a presentation, a press release .... the list goes on.  
Let’s give Google credit where it’s due. They are serving a very useful purpose. But if machine translation ever does become as good as human translation, I dare say it won’t just be the likes of professional translators that will have to worry. No, a machine that actually thinks like a human is a more menacing prospect altogether.
 
A Dutch version of the text is available for readers below. The ‘ass’ responsible for translating it was Google....


Is machine translation allemaal dat het gekraakt maximaal zijn?
Vertalers zijn zwoegende weg met computer-aided translation systemen voor meer dan een decennium lang. Het idee erachter is eigenlijk heel eenvoudig. Tekst vertaald door een vertaler is opgeslagen in een geheugen bank, niet als een complete document, maar als stukken tekst, dat wil zeggen individuele woorden, clausules, zinnen of alinea's. Als een (bijna) identiek stuk tekst verschijnt in een volgende document voor vertaling, zal het CAT-software blijkt de opgeslagen stuk tekst als de meest ideale optie voor de vertaler te gebruiken. Natuurlijk, om efficiënt te kunnen werken, heb je een geheugen bank die is opgebouwd over een langere periode van tijd, met duizenden en duizenden van deze brokken. Sleutel tot de uitvoerbaarheid is de opbouw van de tekst. En terwijl we wenst te noemen machine translation technologie, in feite om doeltreffend te zijn in de eerste plaats, al de ezel werk moet worden gedaan door mensen. Immers, aan het eind van de dag, het is de vertaler die vestigt op de meest passende vertaling. CAT heeft zijn beperkingen: het komt pas echt tot zijn recht wanneer extreem repetitieve teksten moeten worden vertaald. Haar voornaamste doel is om de consistentie te waarborgen, in plaats van een short-cuts nemen.
Google heeft geen traag katoen op dit concept en ziet u waarom. Door het hebben van web-gebaseerde gegevens die beschikbaar zijn in andere talen, informatie is nu op de vingertoppen van een veel groter aantal gebruikers wereldwijd. Om dit te doen, hebben de onderzoekers bij Google zijn trawlvisserij het web en bijpassende duizenden, zo niet miljoenen documenten en webpagina's die al zijn professioneel vertaald (door mensen). Al deze gegevens zijn ingevoerd om een databank te grote hoeveelheden van meertalige teksten te genereren, dit in aanvulling op meer primitieve machine-vertaling methoden (die enkele woorden om te zetten in plaats van zinnen). Al deze gegevens worden vervolgens gebruikt als bron voor hun vertaling gereedschap. Het is moeilijk voor mij als een professionele vertaler te kloppen in hun motieven om dit te doen, immers, informatie beschikbaar stellen voor een breder publiek wereldwijd is een lovenswaardig doel, ook al zijn hun motieven kunnen worden in de eerste plaats commercieel. Ik heb alleen te denken aan de frustratie heb ik geleden probeerde te koppen of staarten van al die Franse, Portugese en Spaanse sites die ik heb gebladerd door de jaren heen te maken. De resultaten taal-wijs kan mooi jammerlijk worden, maar Google verdient een beetje een thumbs-up voor hun inspanningen in het terugschroeven van de grenzen.
Er zijn veel maren hier wel. Als ik probeer een wettelijk contract omzetten van de ene taal naar de andere door machine translation, de kans is, als het is geschreven in Legalese, zal het komen als een mooi vervormd in de doeltaal worden gemaakt en dus vrij nutteloos voor iedereen wensen te lezen en te interpreteren serieus. Alleen een dwaas zou tekenen op de stippellijn en maken het juridisch bindend. En als een klant zou ik verwachten dat alle documentatie of productspecificaties ik ontvangen van een buitenlandse leverancier aan niets minder dan woord-perfect zijn in mijn eigen taal. Zelfs alledaagse volkstaal, gebruikt bijvoorbeeld in films of documentaires, zou mislukken van de test: de inflexibiliteit van gereedschapswerktuigen in het vertalen van uitdrukkingen en spreektaal gebruikt voor ondertitels zou resulteren in een tamelijk bizarre interpretaties. We kunnen geven van veel meer situaties in het echte leven waar machine translation gewoon niet zou werken: een sollicitatie, een gedicht, een product handleiding, een liefdesbrief, een menu, instructies voor de medicatie, een presentatie, een persbericht .. .. de lijst gaat.
Laten we Google-krediet waar het verschuldigd. Ze zijn het bedienen van een zeer nuttig doel. Maar als machine translation heeft ooit zo goed als menselijke vertaling, durf ik zeggen dat het niet alleen het graag van professionele vertalers die zullen moeten zorgen. Nee, een machine die daadwerkelijk denkt als een mens is een meer dreigend vooruitzicht helemaal.

Thursday, 8 October 2009

Baltic bloomer

Imagine my horror a couple of weeks ago when a colleague pointed out that I’d swapped Latvia for Lithuania in a translation job I'd been working on. An easy mistake for a layman maybe, but as a Geographer, I pride myself on making such simple distinctions. No sooner had my embarrassment subsided, than my colleague sent me a link to an article reporting that the publishers of the Bosatlas, regarded as something as a national treasure in Dutch classrooms, had made a similar gaffe. In a promotional give-away, a mini-atlas called the Boskabouter, they too had inadvertently switched the names. My blushes were spared further on discovering that an even more damaging faux-pas had been made by the Czech Republic’s football federation when their national team faced Lithuania in an international friendly last year. Not only were the Latvian team pictured in the match programme, they had the Latvian national anthem played before the game. The mix-up cost two of the federation’s officials their jobs.
The confusion that many people seem to have with the Baltic States has apparently given the Lithuanians some food for thought. Last year, in an effort to raise its profile and attract new investment, a commission led by the prime minister suggested the option of changing the name of the country. Perhaps not without reason: in Lithuanian, the country is called Lietuva. And after all, as a Lithuanian, you wouldn’t want to see foreign capital being channeled into Riga rather than Vilnius on the basis of a mere muddle, would you?
Fortunately, my mistake was picked up before it ever reached the client. I should imagine the lay-out man or woman at Noordhoff Uitgeverij in Groningen will be even redder-faced. And as for the two Czech officials, all I can say is, I have every sympathy - it’s an easy mistake to make!

Wednesday, 23 September 2009

Gezellig vertalen


Sometimes, the life of a freelance translator is like waiting at a bus stop: you wait days for the next job and then, all of a sudden, three assignments turn up at once. Killing time in between jobs is a necessary part of the trade.
Last week, one whole day passed without a phone call or email, so I started musing on a particularly irksome translation I’d done a fortnight ago full of awkward Dutch words. I did the wordsmith’s equivalent of twiddling my thumbs and decided to make a list of the ten Dutch words I find most difficult to translate into English. I came up with the following (in no particular order): inhoudelijk, uitgangspunt, strak, inzichtelijk, structureel, uitwerken, overzichtelijk, afstemmen, vaststellen, toetsen.
Of course, they all have dictionary definitions, but notoriously, dictionaries fail to elaborate on the context. When I first started translating I used to keep a glossary of such difficult words. This way I thought I’d cracked it, but that was until the next time I came across the word and my glossary failed to live up to expectations. In the end I gave up. All I was doing was building my own dictionary and the circle was complete.
In my time as a translator I’ve translated uitgangspunt as ‘point of departure’, ‘basis’, ‘basic tenet’, ‘(underlying) principle’, ‘starting point’, ‘baseline’, ‘(basic) assumption’, ‘[the] idea behind’, ‘objective’ and probably many, many more. The point is it depends on the context. A scientific text will differ in this respect from an administrative text. The register of the text will likewise determine the choice of words. And sometimes, however narrow the context, the translation quite never fits the bill.
So, should overzichtelijk be written in English manageable, well-organised, easy-to-follow or easy-to-understand?
And whereas the literal meaning of structureel is structural, structureel krapte is probably best translated as a chronic shortage. Recently, an agency asked me to consider dropping my rates and I answered, “Dat wil ik liever per opdracht beslissen, ik ga mijn prijzen niet structureel verlagen”. I suppose you would best translate structureel here (an adverb) as ‘across the board’, or ‘as a blanket measure’.
And is a strakke pak a sharp suit or a close-fitting suit? There’s a big difference. I’d say you have to see the suit first. If not, you might as well use your intuition and pick the word that you think the paying customer would most like to see.
It just goes to show that what’s an everyday word in one language is not easily translatable in another.

Continuing to muse, I thought I would enter my top ten words in Google and see what it threw up. In fact, I ended up with over 10,000 documents in Dutch, the vast majority of which were bestuurlijke texts, that is texts written by public bodies such as local authorities, government departments and NGOs. (See, even bestuurlijk is difficult to translate).
- De inhoudelijke uitgangspunten waren beschreven in de uitgangspuntennota. Voor de inzichtelijkheid zijn de hoofdzaken hiervan in dit rapport opgenomen....
- Wat de conceptueel-inhoudelijke benadering betreft, wordt uitgegaan van volgende drie principes...
- Het bereiken van voldoende afstemming qua leerlingenprofiel en abstractieniveau...
- Actualiseren van de opzet en uitwerking van het gemeentelijke besturingsmodel...
It would seem there is a whole army of mandarins churning out this kind of stuff. And, needless to say, that ‘irksome translation’ I’d done two weeks ago was a bestuurlijk document.

A customer should never assume (ervan uitgaan) that translation is a process of converting a source word into the target language. In this respect, English can present a vast array of possibilities each with its own subtlety of meaning, so the decision to use a particular word is often only made after a long and complex thought process. A good translator will not simply pick the first one that is listed in Van Dale.

Sunday, 20 September 2009

Sunday, 16 August 2009

Danke, aber nein danke (part II)

Has Alemannia got wind of the discontent felt among Roda fans (see below)? No sooner had I posted my blog last week than these posters (in Dutch) started appearing all over Kerkrade and Heerlen, advertising Alemannia Aachen's inaugural match of the 2009/10 season at the new Tivoli stadium in the city (Alemannia currently plays in the 2nd division of the German Bundesliga).
A number of years ago, Alemannia qualified for the UEFA cup, but because of stadium requirements at the time, they were not allowed to play at the old Tivoli. In consequence, the club submitted a request to UEFA to have their home games played at the new purpose-built, 20,000 seater, Parkstad Limburg stadium, Roda JC's home ground - just 12 kilometres away. A logical choice it would seem. However, the request was turned down on the grounds that clubs representing their football association were not permitted to play outside the territory in which the association's jurisdication held sway. Instead, they played their home legs at FC Cologne's ground, over 60 kilometres away.
Talk about daft!

Friday, 7 August 2009

Danke, aber nein danke

Roda JC, the Kerkrade-based football club which plays in the top-flight of the Dutch league, this week issued a statement advising its fans to desist from singing German-language chants at its home matches. The club has been moved to issue the statement following the findings of a survey carried out by Club Positioning Matrix, which reveal that the club has a 'weinig sympathiek Duits imago'. The reason is primarily financial, the club say. Income from television rights is distributed on the basis of a club's positioning in the matrix and Roda come off pretty badly in the rankings, supposedly as a result of this 'German' image. More than anything, the management wants to develop its image as a mainstream 'Dutch' club (thus enhancing its chances of increased revenue) and ditch its regional identity: the insinuation is that a German image is bad for the club.
Traditionally, goals scored by the home club at Parkstad Stadion Limburg are celebrated in the German style, with the tannoy system booming out 'Danke', to be reciprocated by a 'Bitte' from those in the stands. Another favourite chant of the crowd is 'Viva Colonia', a song written by Cologne-based band De Höhner, popular at Carnaval as well as on the terraces in the Rheinland region of Germany (and in Dutch Limburg).
In a country where anti-German sentiment is often simmering just below the surface, perhaps one shouldn't be surprised to see the words 'weinig sympathiek' (less favourable) and 'Duits' (German) used in the same breath. Nevertheless, I find the juxtaposition of words pretty woeful.
Equally lamentable are the attempts by the club management to snuff out regional identity. Sadly, many Limburgers themselves are only too oblivious to the cultural and linguistic roots they share with their neighbours. After all, it is only by a quirk of fate that 'Limburg' became part of the Netherlands. If history had taken a different turn, Limburgers might well have become Germans or Flemings. The 'vernederlandsing' of Limburg has been a slow, gradual process.
Linguistically, the Limburg dialect is Ripuarian in origin, spoken widely in the region to the west of the Rhine. In the south-eastern corner of Limburg, German for a long time competed with Dutch as the Lingua Franca. Up until the 1930's, German was still being used in some churches. In Kerkrade, up until 1911, there was a German language newspaper (see above). In fact, culturally speaking, there's a whole raft of customs and traditions which connect Limburg more closely with the Rheinland than with 'Holland': food architecture, industry (mining), religion, culture (Carnaval, the schutterijen, brass bands, Schlagermusik, etc.)
The Roda management may do its best to stamp out the German chanting, but old habits die hard and thankfully traditional regional affinities persist, if only sometimes in a latent, subliminal form. (Ironically, the name Roda refers to the region known as the Land of Roda which straddles the border between Kerkrade in the Netherlands and Herzogenrath in Germany.)
As the addage goes, you can take the Kerkradenaar out of Kerkrade, but you can't take Kerkrade out of the Kerkradenaar.

Monday, 27 July 2009

Woldgate

A few days ago, I was driving along Woldgate in the East Riding of Yorkshire, in countryside not dissimilar to South Limburg, after spending a terrific afternoon at Flamborough Head. This C-classified road is as straight as an arrow and is marked on the Ordnance Survey map as a Roman Road, which supposedly linked York with the coast from where I'd just come. There are 'gates' aplenty in the North-East of England, most notably in York (Stonegate, Goodramgate and Skeldergate). These 'gates' have nothing to do with the gateways that are dotted around the ancient city walls of York, which are known locally as bars (Bootham Bar, Monk Bar). The names of these streets originate from the Viking word 'gata', meaning 'road' or 'way'. Presumably, the name Woldgate is derived from the same origin.
We know that much of the region was under Viking domination prior to 950, where Danish law and custom were observed, hence the name Danelaw to denote the territory that split England in two along a line from the Dee to the Thames. To the south-west of the line, Anglo-Saxon law held sway.
The Vikings first invaded England in around 800 and over the next 50 years or so seized large amounts of territory in Northern England including the Northumbrian (Anglo-Saxon) capital of Eoforwik, which was renamed Jorvik (now York). Viking settlements were established in the fertile regions in the North and East Anglia (alongside existing Anglo-Saxon ones). The Danes divided Yorkshire into three parts called 'Ridings' (Old Norse for 'third') for administrative purposes and these three regions were to survive for many centuries.
It was not far from here, at Stamford Bridge, that in 1066 the Vikings suffered their final defeat on English soil at the hands of King Harold, who himself lost a more famous battle two weeks later at Hastings.
The Yorkshire Wolds abounds with strange-sounding place names: Wetwang, Thwing, Langtoft, Kirkby Grindalythe, Weaverthorpe, Ruston Parva, Fangfoss and Nunburnholme to name but a few. The exact origins of these place names is uncertain, but what is for sure is that they include common Viking elements: Toft (a plot of land or farm), Thorpe (a small farm, hamlet of outlying settlement), By (a farm or village), Holm (water meadow) and Foss (a ditch). It is thought that 40 percent of place names in the East Riding of Yorkshire owe their existence to a Scandinavian presence, even though these settlements might have already been in existence during Saxon rule. And Viking and Saxon settlements would have coexisted with each other.
Some language historians have theorised that the mix of the two languages was responsible for kick-starting the 'simplification' of Old English into Middle English. Until the arrival of the Scandinavians, Old English was a strongly inflected tongue (like modern German) where common words relied on word-endings to convey a meaning for which we now use prepositions, like 'to', 'with' and 'from'. And, for example, by adding an -s, plurals were made less complicated. Only a few old noun inflections have survived, such as geese, mice and children. The complications of the language, it is said, were gradually ironed out as a way for Anglo-Saxon and Scandinavian speakers to understand each other better, through an ongoing process of 'pidginsation'. Of course, this shift didn't take place overnight, but over a period of centuries.
In addition to their place names, the Vikings donated many words to the English language (common words like get, hit, leg, low, root, skin, same, want and wrong are all of Scandinavian origin) and even more words survive in dialect. The word 'laik', for example, still enjoys popular use by children throughout the North-East and Cumbria, where it means 'to play'. A famous Danish product that many people will have played with in their childhood is Lego, which comes from the Danish 'leg godt' meaning to 'play well'.