Python Multithreading
Multithreading is similar to executing multiple different programs at the same time. Running multithreads has the following advantages:
- Using threads can put tasks that occupy a long time in a program into the background for processing.
- The user interface can be more attractive. For example, when a user clicks a button to trigger the processing of certain events, a progress bar can pop up to show the processing progress.
- The running speed of the program may be accelerated.
- Threads are more useful for implementing some waiting tasks such as user input, file reading and writing, and network data sending/receiving. In such cases, we can release some precious resources such as memory usage.
Threads still differ from processes during execution. Each independent process has an entry point for program execution, a sequential execution sequence, and an exit for the program. However, threads cannot execute independently; they must depend on an application, which provides multiple thread execution controls.
Each thread has its own set of CPU registers, called the thread's context. This context reflects the state of the CPU registers when the thread last ran.
The instruction pointer and stack pointer registers are the two most important registers in the thread context. A thread always runs in the context of a process, and these addresses are used to identify memory in the address space of the process that owns the thread.
- Threads can be preempted (interrupted).
- While other threads are running, a thread can be temporarily suspended (also called sleeping) -- this is thread yielding.
Start Learning Python Threads
There are two ways to use threads in Python: with functions or by using a class to wrap thread objects.
Functional approach: call the start_new_thread() function in the thread module to create a new thread. The syntax is as follows:
thread.start_new_thread ( function, args[, kwargs] )
Parameter description:
- function - The thread function.
- args - The arguments passed to the thread function; it must be a tuple type.
- kwargs - Optional arguments.
Example (Python 2.0+)
The output of running the above program is as follows:
Thread-1: Thu Jan 22 15:42:17 2009 Thread-1: Thu Jan 22 15:42:19 2009 Thread-2: Thu Jan 22 15:42:19 2009 Thread-1: Thu Jan 22 15:42:21 2009 Thread-2: Thu Jan 22 15:42:23 2009 Thread-1: Thu Jan 22 15:42:23 2009 Thread-1: Thu Jan 22 15:42:25 2009 Thread-2: Thu Jan 22 15:42:27 2009 Thread-2: Thu Jan 22 15:42:31 2009 Thread-2: Thu Jan 22 15:42:35 2009
The termination of a thread generally relies on the natural end of the thread function; you can also call thread.exit() in the thread function, which raises a SystemExit exception to achieve the purpose of exiting the thread.
Thread Module
Python provides support for threads through two standard libraries, thread and threading. thread provides low-level, primitive threads and a simple lock.
Other methods provided by the threading module:
- threading.currentThread(): Returns the current thread variable.
- threading.enumerate(): Returns a list containing running threads. Running means after the thread starts and before it ends, not including threads before start and after termination.
- threading.activeCount(): Returns the number of running threads, same result as len(threading.enumerate()).
In addition to using methods, the thread module also provides the Thread class to handle threads. The Thread class provides the following methods:
- run():Method used to represent thread activity.
- start():Start thread activity.
- join([time]):Wait until the thread terminates. This blocks the calling thread until the thread's join() method is called to terminate - normal exit or an unhandled exception - or an optional timeout occurs.
- isAlive():Return whether the thread is active.
- getName():Return the thread name.
- setName():Set the thread name.
Creating Threads with the Threading Module
To create a thread using the Threading module, directly inherit from threading.Thread, and then override the __init__ method and the run method:
Example (Python 2.0+)
The result of running the above program is as follows:
Starting Thread-1 Starting Thread-2 Exiting Main Thread Thread-1: Thu Mar 21 09:10:03 2013 Thread-1: Thu Mar 21 09:10:04 2013 Thread-2: Thu Mar 21 09:10:04 2013 Thread-1: Thu Mar 21 09:10:05 2013 Thread-1: Thu Mar 21 09:10:06 2013 Thread-2: Thu Mar 21 09:10:06 2013 Thread-1: Thu Mar 21 09:10:07 2013 Exiting Thread-1 Thread-2: Thu Mar 21 09:10:08 2013 Thread-2: Thu Mar 21 09:10:10 2013 Thread-2: Thu Mar 21 09:10:12 2013 Exiting Thread-2
Thread Synchronization
If multiple threads modify a certain piece of data together, unpredictable results may occur. To ensure data correctness, multiple threads need to be synchronized.
Using the Lock and RLock of Thread objects can implement simple thread synchronization. These two objects both have acquire and release methods. For data that requires only one thread to operate at a time, its operations can be placed between the acquire and release methods. As follows:
The advantage of multithreading is that it can run multiple tasks at the same time (at least it feels that way). However, when threads need to share data, there may be data desynchronization issues.
Consider such a situation: all elements in a list are 0. Thread "set" changes all elements to 1 from back to front, while thread "print" is responsible for reading the list from front to back and printing.
Then, it is possible that when thread "set" starts changing, thread "print" comes to print the list, and the output becomes half 0s and half 1s. This is data desynchronization. To avoid this situation, the concept of locks is introduced.
A lock has two states - locked and unlocked. Whenever a thread such as "set" wants to access shared data, it must first acquire the lock; if another thread such as "print" has already acquired the lock, then let thread "set" pause, which is synchronous blocking; after thread "print" finishes accessing and releases the lock, let thread "set" continue.
With such handling, when printing the list, either all 0s are output or all 1s are output, and the awkward scene of half 0s and half 1s will no longer occur.
Example (Python 2.0+)
Thread Priority Queue
The Queue module in Python provides synchronized, thread-safe queue classes, including FIFO (first-in-first-out) queue Queue, LIFO (last-in-first-out) queue LifoQueue, and priority queue PriorityQueue. These queues all implement lock primitives and can be used directly in multithreading. Queues can be used to achieve synchronization between threads.
Common methods in the Queue module:
- Queue.qsize() Returns the size of the queue
- Queue.empty() If the queue is empty, return True, otherwise False
- Queue.full() If the queue is full, return True, otherwise False
- Queue.full corresponds to the maxsize size
- Queue.get([block[, timeout]]) Gets the queue, timeout is the waiting time
- Queue.get_nowait() Equivalent to Queue.get(False)
- Queue.put(item, block=True, timeout=None) Writes to the queue, timeout is the waiting time
- Queue.put_nowait(item) Equivalent to Queue.put(item, False)
- Queue.task_done() After completing a piece of work, the Queue.task_done() function sends a signal to the queue that the task has been completed.
- Queue.join() Actually means to wait until the queue is empty before performing other operations.
Example (Python 2.0+)
The result of running the above program:
Starting Thread-1 Starting Thread-2 Starting Thread-3 Thread-1 processing One Thread-2 processing Two Thread-3 processing Three Thread-1 processing Four Thread-2 processing Five Exiting Thread-3 Exiting Thread-1 Exiting Thread-2 Exiting Main ThreadOther Extensions