Показаны сообщения с ярлыком in English. Показать все сообщения
Показаны сообщения с ярлыком in English. Показать все сообщения

вторник, 17 марта 2020 г.

Running Anki from a Flash Drive

if you have made a destructive change on one computer and have an undamaged copy on another computer, you may wish to start Anki without syncing in order to use the full sync option without first downloading the changes. Similarly, if you are experiencing problems with Anki, you might want to (or might be instructed to) disable add-ons temporarily to see if one might be causing the problem. You can do both of these things by holding down the Shift key while starting Anki.
It is possible to specify a custom folder location during startup. This is an advanced feature that is primarily intended to be used with portable installations, and we recommend you use the default location in most circumstances.
The syntax to specify an alternate folder is as follows:
anki -b /path/to/anki/folder
  • If you have multiple profiles, you can pass -p <name> to load a specific profile.
  • To change the interface language, use -l <iso 639-1 language code>, such as “-l ja” for Japanese.
If you always want to use a custom folder location, you can modify your shortcut to Anki. On Windows, right-click on the shortcut, choose Properties, select the Shortcut tab, and add “-b \path\to\data\folder” after the path to the program, which should leave you with something like
"C:\Program Files\Anki\anki.exe" -b "C:\AnkiDataFolder"
You can also use this technique with the -l option to easily use Anki in different languages.
On Windows, you should use a backslash (\) not a forward slash (/).
On a Mac there is no easy way to alter the behaviour when clicking on the Anki icon, but it is possibile to start Anki with a custom base folder from a terminal:
open /Applications/Anki.app --args -b ~/myankifolder

Running from a Flash Drive

On Windows, Anki can be installed on a USB / flash drive and run as a portable application. The following example assumes your USB drive is drive G.
  • Copy the \Program Files\Anki folder to the flash drive, so you have a folder like G:\Anki.
  • Create a text file called G:\anki.bat with the following text:
g:\anki\anki.exe -b g:\ankidata
If you would like to prevent the black command prompt window from remaining open, you can instead use:
start /b g:\anki\anki.exe -b g:\ankidata
  • Double-clicking on anki.bat should start Anki with the user data stored in G:\ankidata.
The full path including drive letter is required - if you try using \anki\anki.exe instead you will find syncing stops working.
Media syncing with AnkiWeb may not work if your flash drive is formatted as FAT32. Please format the drive as NTFS to ensure media syncs correctly.

среда, 15 января 2020 г.

How to Transcribe Audio Notes Easily Using YouTube Auto Captions

YouTubeYou might sometimes need a quick transcription of some audio notes and you might not feel like rewinding back and forth. Luckily, there is a very easy way to do it using YouTube’s auto-captioning system.

The results aren’t always perfectly accurate. However, all you need to do is go through the text a little and correct the mistakes. This will still end up saving you the time you’d have spent doing it some other way.

Here’s how you can can use YouTube to transcribe your audio notes.

USING YOUTUBE’S AUTO CAPTION SYSTEM TO TRANSCRIBE TEXT

The process of using Auto Caption to transcribe a text is quite simple.

Basically, all you need is the audio file you want to transcribe, an image file (it doesn’t matter what it shows) and a YouTube account. The rest of the tools are easily found online.

First, you’ll need your audio file to become a video (at least, in YouTube’s eyes). This is done just the way you would upload a song to YouTube (if you want to convert a YouTube video to MP3, that’s easily done, too).

Step 1: The easiest way to do this is to use a service like TunesToTube, which will turn your audio file into a YouTube video. So, go to TunesToTube and connect your YouTube account to the service. If you’re not logged into YouTube, you’ll need to login.

YouTube auto caption transcribe

Step 2: Give TunesToTube the permission to manage your YouTube account.

YouTube auto caption transcribe permission

Step 3: Your next move is to add your audio and image files. This is done via TunesToTube’s simple interface – click Upload Files. Add whatever title and description you want and you don’t have to add any tags if you don’t want to.

YouTube auto caption transcribe  upload

NOTE: The only important thing is for the audio to be in a language that YouTube’s auto caption feature understands. Those languages are English, Dutch, French, German, Italian, Japanese, Korean, Portugese, Russian and Spanish.

Step 3: Once you’ve browsed for your files and have added them to TunesToTube, don’t forget to fill in the CAPTCHA. After that, you can go ahead and click Create Video!

YouTube auto caption transcribe youtube

Step 4: Wait for your video to be uploaded. The time it will take to upload depends on how large your audio file and image are, how fast your internet connection is and, last but not least, the load on YouTube’s servers at that particular moment.

After that, wait for your video to process – you can see whether the process is over in YouTube’s Video Manager if you’re getting impatient.

Step 5: When the video’s processing has been completed, know that it will still take a while before it is captioned. The duration depends again on YouTube’ server load at that time of the day. I’ve tried doing this with a few videos, all of them a few minutes long. In some cases it took just a few minutes, while there was a time lag when the process lasted 15 minutes. It’s still quick, though.

When it’s done, you’ll see an icon (like the one shown in the screenshot) marked below your video, called Transcription. Click the Transcription button.

YouTube auto caption transcribe transcription done

Step 6: You can now copy your transcription and correct the mistakes. If the person speaking does so in a clear fashion and there’s not too much noise around, it should be pretty accurate.

YouTube auto caption transcribe

COOL TIP: If you don’t understand what YouTube has transcribed, you can just click that line and you’ll be taken to the specific moment in the video, so you don’t have to browse through all of it to find something.

CONCLUSION

If you need a quick transcription of an audio memo, this is an easy way to do it. The results are not always perfect, but it can be helpful.


вторник, 17 сентября 2019 г.

Bulk Generating Cloze Deletions for Learning a Language with Anki

One challenge in learning a second language is the sheer amount of vocabulary required for fluency. For example, one study has shown that the average eight year old knows 8000 words, while the average adult might know somewhere between 20,000 and 35,000 words. University graduates might know upwards of 50,000 words depending on their specialization.







As an adult learning a second language, exposure to a lot of different words in their context is one of the best ways to learn vocabulary. A common test for learning words in context is called a cloze test.
“A cloze test (also cloze deletion test) is an exercise, test, or assessment consisting of a portion of text with certain words removed (cloze text), where the participant is asked to replace the missing words. Cloze tests require the ability to understand context and vocabulary in order to identify the correct words or type of words that belong in the deleted passages of a text. This exercise is commonly administered for the assessment of native and second language learning and instruction.”
— Wikipedia





For example, here is a cloze test from the previously linked Wikipedia article.
Today, I went to the ________ and bought some milk and eggs. I knew it was going to rain, but I forgot to take my ________, and ended up getting wet on the way.
The interesting thing about cloze tests is that they force you to understand both vocabulary and the context where it is used. In the previous example, the first blank is preceded by the word the, and therefore must be followed by a noun, adjective, or adverb. However, a conjunction follows the blank; the sentence would not be grammatically correct if anything other than a noun were in the blank. Correctly completing the test forces the learner to understand both vocabulary and grammatical rules of vocabulary usage. This makes cloze tests ideal for learning and testing a second language.
I use Anki — a program that helps make remembering things as easy as possible — to automate how frequently I complete cloze tests. Because Anki is a lot more efficient than traditional study methods, you can greatly decrease your time spent studying to learn vocabulary. Unfortunately, Anki requires you to create your own cloze tests, which can be a tedious and time consuming process. Being a software engineer, I began looking at how to automate the process of creating cloze tests for Anki and came up with an automatic method using publicly available data.
The rest of this article describes the process I used for automatically creating cloze tests for import into Anki, followed by links to download pre-packaged Anki decks for learning French as an English speaker created using this method..

Raw Sentences

The first step is to collect a lot of sentences in your target language. Ideally, sentences would have a translation in your native language to help understand any contextual clues for completing the cloze test. The website Tatoeba provides a collection of user generated sentences and translations, which makes an ideal data set for sentences. All data is available for download under a Creative Commons Attribution 2.0 license (CC-BY 2.0) as a set of simple tab separated files. The files come in two formats: first, a list of all sentences and the language they are written in, and second, a list of links between the sentences that shows which sentences are translations of each other.
The sentences dataset looks like the following, where the first number is the sentence identifier, followed by the language, followed by the text.
107 deu Ich kann mich nur fragen, ob es für alle anderen dasselbe ist.
1115 fra Lorsqu'il a demandé qui avait cassé la fenêtre, tous les garçons ont pris un air innocent.
6557647 eng I sing of arms and the man, made a fugitive by fate, who first came from the coasts of Troy, to Italy and the Lavinian shores.
The links dataset provides a simple mapping between sentence numbers, showing which sentences are translations of each other.
1 77
1 1276
1 2481
1 5350
1 5972
1 180624
1 344899
1 345549
1 380381
1 387119

Frequency Lists

The sentences from Tatoeba provide a great starting point for generating cloze deletions. The question still remains, which word to use for the cloze? That is, which word will be the blank in the fill in the blank problem?
Not all words are created equal, some are so common that they are used in almost every sentence and are therefore not that interesting to learn, others are used so infrequently that they would not be useful for everyday speech. In the middle are the 20,000 or so words that are used regularly enough that they can provide a basis for fluency in a second language.
A sorting of words into the frequency with which they appear in a corpus of language is called a frequency list. Wiktionary provides a list of frequency lists available in many languages. I used a premade list made available under an MIT license on Github. This list tracks the 50,000 most frequently used words in TV and movie subtitles maintained by the OpenSubtitles project.
I chose the cloze deletion to test using the same method used by clozemaster.
The cloze deletion to test, or the blank in the sentence, is the least common word in the sentence within the [5]0,000 … most common words in the language. In other words, for a given sentence all the words in the sentence are checked against the top [5]0,000 words in a fequency list for that language. The least common word is then used as the cloze test. In this way the vocab learned via clozemaster is the most difficult of the most common.

Generating the Cloze Tests

Using the original sentences with translations, and a frequency list, I created a set of cloze deletions using a Python script.

Finding target and native language sentences

The first step was to extract out the sentences from Tatoeba that are available in my native language and target language. I used a simple grep expression to create two files from the original sentence data.
grep -E '\teng\t' sentences.csv > native_sentences.csv
grep -E '\tfra\t' sentences.csv > target_sentences.csv

Choosing the cloze word

The following function takes a sentence and a frequency list (as a map), and chooses a cloze word. I skipped any words that were capitalized and removed any words that were two characters or shorter. Any words left over were checked against the frequency list and the minimum frequency word was used for the cloze. If there were no words in the sentence that had frequency data attached, a random word was chosen.
def find_cloze(sentence, frequency_list):
    """
    Return the least frequently used word in the sentence by.
    If no word is found in the frequency list, return a random word.
    If no acceptable word is available (even random), return None
    """
    # Remove punctuation
    translator = str.maketrans(string.punctuation, ' '*len(string.punctuation))
    sentence = sentence.translate(translator)

    max_frequency = 50001  # Frequency list has 50,000 entries
    min_frequency = max_frequency

    min_word = None
    valid_words = []
    for word in sentence.split():
        if word.isupper() or word.istitle():
            continue  # Skip proper nouns
        if len(word) <= 2:
            continue  # Skip tiny words

        valid_words.append(word)

        word_frequency = int(frequency_list.get(word.lower(), max_frequency))
        if word_frequency < min_frequency:
            min_word = word
            min_frequency = word_frequency

    if min_word:
        return min_word
    else:
        if valid_words:
            return random.choice(valid_words)
        else:
            return None

Synthesizing speech

I used Amazon Polly to attach audio samples to the target language. The script randomly chooses a voice in the target language to add a bit of variety to the audio.
import boto3

polly = boto3.client('polly')

def synthesize_speech(text, filename):
    """
    Synthesize speech using Amazon polly
    """
    voices = ['Celine', 'Mathieu', 'Chantal']
    voice = random.choice(voices)

    response = polly.synthesize_speech(
        OutputFormat='mp3',
        Text=text,
        VoiceId=voice
    )

    output = os.path.join("./out/", filename)

    if "AudioStream" in response:
        with closing(response["AudioStream"]) as stream:
            with open(output, "wb") as outfile:
                outfile.write(stream.read())

Generating a CSV for Anki

Given the algorithm for choosing a cloze word, and a method for synthesizing speech, the last step is to generate a CSV file from the data sets that can be input into Anki. The CSV generation requires a file of french and english sentences, the links file between the translations, and the frequency list. It matches sentences with their translations, chooses a cloze deletion, and synthesizes an audio sample of the sentence.
def make_index(path, delimiter, value=1):
    """
    Given a CSV reader, return a map between the first column and the
    column specified by value.
    """
    d = dict()
    with open(path, newline='') as f:
        reader = csv.reader(f, delimiter=delimiter)
        for row in reader:
            d[row[0]] = row[value]
    return d


def generate(french_sentence_file,
             english_sentence_file,
             links_file,
             frequency_list_file):
    # Make index between sentence number and rest of csv
    print("Making indexes ...")

    french = make_index(french_sentence_file, '\t', value=2)
    english = make_index(english_sentence_file, '\t', value=2)
    links = make_index(links_file, '\t')

    # Make index between word and usage frequency
    frequency = make_index(frequency_list_file, ' ')

    print("Generating clozes ...")
    with open("out.csv", 'w', newline='') as outfile:
        writer = csv.writer(outfile, delimiter='\t',
                            quotechar='|', quoting=csv.QUOTE_MINIMAL)

        # For each French sentence
        for fra_number, fra_sentence in french.items():
            # Lookup English translation
            eng_number = links.get(fra_number)
            if not eng_number:
                continue  # If no English translation, skip

            eng_sentence = english.get(eng_number)
            if not eng_sentence:
                continue  # If no English translation, skip

            # Find the cloze word
            fra_cloze_word = find_cloze(fra_sentence, frequency)
            if not fra_cloze_word:
                continue  # If no cloze word, skip

            clozed = fra_sentence.replace(fra_cloze_word,
                                          '{{{{c1::{}}}}}'.format(fra_cloze_word)

            # Generate audio
            audio_filename = 'fra-{}-audio.mp3'.format(fra_number)
            if not os.path.isfile(audio_filename):
                synthesize_speech(fra_sentence, audio_filename)

            writer.writerow([fra_number,
                             clozed,
                             eng_number,
                             eng_sentence,
                             '[sound:{}]'.format(audio_filename)])

    print("Done.")

Importing into Anki

Given the CSV file and a list of mp3 audio samples, the next step is to import the data into Anki. The audio is imported by copying the files into Anki’s collection.media folder. See the Anki manual for more information.
Once the media files are imported, you can import the CSV file generated by the cloze script into Anki. To import a file, click the File menu and then “Import”. For more information, see the Anki manual.

Downloads

I used this process to generate French cloze deletions suitable for English speakers learning French. You can download pre-made packages that you can import into Anki to start learning French here.

воскресенье, 21 апреля 2019 г.

successful movement of ankidroid folder to external SDcard


in my phone I have only 16gig internal memory and about 3000 cards that have media (images, sound, short video clips with mp4 formats) so the phone's memory was not enough. I did this to move ankidroid folder to external memory.

1- close the app, go to to this address on external memory: android/data 
3- make a new folder named: com.ichi2.anki
4- move the ankidroid folder (on internal memory) to the above mentioned address so it would be:  /storage/sdcard1/android/data/com.ichi2.anki/ankidroid
5-open the ankidroid app and go to settings/advanced and set the following address for data location instead of storage/ankidroid:
 /storage/sdcard1/android/data/com.ichi2.anki
6-enjoy
maybe in your phone the beginning of the address changes. for exact address you can retrieve the address for example by a sample filej's properties (or details) using x plore or es file manager



edition of line 5: /storage/sdcard1/android/data/com.ichi2.anki/ankidroid
filej's→ file manager's

Thanks so much for showing us how to do this.
Works great! As suggested, I also needed to change the beginning of the path in my case /storage/extSdCard/...

13.05.17real...@gmail.com
your welcome


On Thursday, April 27, 2017 at 4:04:53 PM UTC+4:30, Thang La wrote:
Thanks so much for showing us how to do this.
Works great! As suggested, I also needed to change the beginning of the path in my case /storage/extSdCard/...

05.10.17pavel...@gmail.com
Thank you. For me this is function only if AnkiDroid app is installed in inner memory of the phone. When I move the app to the SD Card, AnkiDroid stops working. Pavel

06.10.17franciscojer...@gmail.com
Thanks, on mine I used the ES File Explorer to pick up the correct patch for my Galaxy Tab E
/storage/extSdCard/Android/data/com.ichi2.anki/Ankidroid

суббота, 13 апреля 2019 г.

PUBLIC Stack Overflow Tags Users Jobs Teams Q&A for work Learn More How to make queries on tatoeba sentence.csv using links.csv in mysql

I want to make an offline sentence dictionary database in my localhost from tatoeba.org dump files in the link tatoebaarchive. As you can notice from the codes, sentences.csv contains language idssuch as eng, fra, cmn, tur and the sentences in that language; and links.csv maps sentences ids to translated parallel sentences.
I want to search a word in a language, say eng, and list pairs of searched word's sentences and translated sentences of it. For example I search "beauty" in English (eng) and list them with its parallel sentences in French (fra).
`sentences.csv` has three columns: `id` `language`  `text`
`links.csv` has two columns: `sentenceId`   `translatedId`
I can make a parallel sentence in three steps.
1) SELECT id FROM sentences WHERE text LIKE('%she is a rare beauty_%') AND language='eng'
       +-------+
       | id    |
       +-------+
       | 21687 |
       +-------+
2) SELECT translatedId FROM links WHERE sentenceId='21687'
      +--------------+
      | translatedId |
      +--------------+
      |       184559 |
      |       517365 |
      |       550067 |
      |      2238371 |
      |      2238372 |
      +--------------+
3) SELECT text FROM sentences WHERE (id ='550067' AND language='fra') OR id ='21687'
      +------------------------------------------+
      | text                                     |
      +------------------------------------------+
      | It is true she is a rare beauty.         |
      | C'est vrai, elle est d'une rare beauté.  |
      +------------------------------------------+
How can I merge three queries into one liner query to get the third step result?
  • Sorry - what exactly is the question? How to import the data to MySQL tables? How to optimize the query speed? How to perform LIKE searches? Your post is very, very unclear... – ılǝ Oct 9 '13 at 7:21
  • Thanks for responding. Core of my question is how to make a query to list parallel sentences that a particular word exist. Like 'It is true she is a rare beauty. > C'est vrai, elle est d'une rare beauté.' – kenn Oct 9 '13 at 7:36
  • how to make a list parallel sentences that a particular word exist - your query is correct. Sorry - it remains unclear if you are asking how to optimize the speed of the query, or how to do a join on the tables, or how to deal with special characters or, or, or – ılǝ Oct 9 '13 at 7:41
  • Example of jukuu.com bilingual sentence searcher exatcly what I want to do. How can I achieve it with sql queries on sentences.csv database? – kenn Oct 9 '13 at 15:51
  • Clarifying my question,I must say it's confusing and it requires subqueries besides my English is not perfect. I want to query a word (say "good") in sentence table from "text" column that contains "eng" language ; and using each found sentence's "id" to make another query from links table, so it will give me "translatedId"s of found sentences and finally using "translatedId"s to make another query in sentences table from "id" column that contains "fra" language. It will list English sentences that includes "good" word with their translated counterparts in French. – kenn Oct 11 '13 at 9:21
  SELECT `sentences`.* FROM 
  `sentences` JOIN 
  `links` ON `id` = `translatedId` 
  WHERE `sentenceId` = (SELECT id FROM sentences WHERE text LIKE('%she is a rare beauty_%') AND language='eng' LIMIT 1);
The result is
  +---------+----------+-------------------------------------------------+
  | id      | language | text                                            |
  +---------+----------+-------------------------------------------------+
  |  184559 | jpn      | 確かに彼女は絶世の美人です。                    |
  |  517365 | deu      | Es ist wahr, sie ist eine seltene Schönheit.    |
  |  550067 | fra      | C'est vrai, elle est d'une rare beauté.         |
  | 2238371 | ber      | S tidet, drusit tsednan ay icebḥen am nettat.   |
  | 2238372 | ber      | Ccbaḥa-nnes drus tin ay tt-yesɛan.              |
  +---------+----------+-------------------------------------------------+