Monday, October 27, 2008

An Advanced Least-Significant-Bit Embedding Scheme for Steganographic Encoding

Authors: Yeuan-Kuen Lee, Graeme Bell, Shih-Yu Huang, Ran-Zan Wang and Shyong Jian Shyu

The 3rd Pacific-Rim Symposium on Image and Video Technology ( PSIVT 2009 )
Tokyo, Japan, January 13th - 16th, 2009
Official Website: http://psivt2009.nii.ac.jp/

Abstract

The advantages of Least-Significant-Bit (LSB) steganographic data embedding are that it is simple to understand, easy to implement, and it results in stego-images that contain hidden data yet appear to be of high visual fidelity. However, it can be shown that under certain conditions, LSB embedding is not secure at all. The fatal drawback of LSB embedding is the existence of detectable artifacts in the form of pairs of values (PoVs). The goals of this paper are to present a theoretic analysis of PoVs and to propose an advanced LSB embedding scheme that possesses the advantages of LSB embedding suggested above, but which also provides an additional level of communication security. The proposed scheme breaks the regular pattern of PoVs in the histogram domain, increasing the difficulty of steganalysis and thereby raising the level of security. The experimental results show that both the Chi-square index and RS index are less than 0.1, i.e., the hidden message is undetectable by the well-known Chi-square and RS steganalysis attacks.



這篇就是我們即將在 PSIVT 2009 發表的論文, 其實內容就是 95 學年度 國科會計畫 的結案報告改寫成論文發表。

國科會計畫編號: NSC 95-2221-E-130-014
最低位元嵌入法的修正模型與安全分析
A Modified LSB Embedding Scheme of Steganography and its Security Analysis
執行期間: 2006/08/01 ~ 2007/10/31
計劃書中文摘要下載:

Comments of Reviewer 1

SUMMARY AND CONTRIBUTIONS: This paper proposed an improved LSB steganographic method. The contribution is that both the Chi-square and RS steganalysis attacks can be resisted.
OVERALL EVALUATION: 7 (strong accept)

COMMENTS ON OVERALL EVALUATION: LSB-basd embedding methods seem to be impractical because an image is usually compressed before transmission.
ORIGINALITY: 3 (moderately original)
REFERENCE TO PRIOR WORK: 4 (excellent reference to prior work)
RELEVANCE/IMPORTANCE TO PSIVT: 3 (of sufficient interest)
CLARITY OF PRESENTATION: 3 (is clear enough)
TECHNICAL CORRECTNESS: 3 (probably correct)
EXPERIMENTAL EVALUATION: 4 (sufficient evaluation or theoretical paper)

Comments of Reviewer 2

SUMMARY AND CONTRIBUTIONS: The author(s) of this paper touch(es) on the simplicity of least significant bit (LSB) embedding and highlight(s) its weakness in the form of pairs of values (PoVs) as steganographic encoding artifacts. A new technique using pseudorandom number generator (PRNG) is employed in an algorithm to modify the method of embedding the secret message bits in the LSB of the target image. This method breaks the correlation between the frequency of these pairs of values commonly caused by LSB embedding. The result of the paper is promising and shows resistance to both the Chi-square and RS steganalysis attacks.

OVERALL EVALUATION:
7 (strong accept)

COMMENTS ON OVERALL EVALUATION: This paper establishes a good model for analysing the effect of PoVs and ventures from there to find a method to avoid the pitfalls of LSB embedding by captilising on the property of pseudorandom number generator. The proposed algorithm effectively prevents successful attacks from both Chi-square and RS steganalysis.

ORIGINALITY:
4 (very original)
REFERENCE TO PRIOR WORK: 3 (references adequate)
RELEVANCE/IMPORTANCE TO PSIVT: 3 (of sufficient interest)
CLARITY OF PRESENTATION: 4 (easy to read)
TECHNICAL CORRECTNESS: 3 (probably correct)
EXPERIMENTAL EVALUATION: 4 (sufficient evaluation or theoretical paper)

Comments of Reviewer 3

SUMMARY AND CONTRIBUTIONS: The Least-significant-bit embedding method is well-known technique in data embedding field, however, this paper proposes an advanced LSB embedding method to improve the lack of traditional LSB embedded method. And the experimental results are enough to verify the goals. The paper is esay to read and understand.

OVERALL EVALUATION:
6 (accept)

COMMENTS ON OVERALL EVALUATION: This system is valuable to data embeddubg scheme.

ORIGINALITY: 3 (moderately original)
REFERENCE TO PRIOR WORK: 3 (excellent reference to prior work)
RELEVANCE/IMPORTANCE TO PSIVT: 3 (of sufficient interest)
CLARITY OF PRESENTATION: 4 (references adequate)
TECHNICAL CORRECTNESS: 3 (probably correct)
EXPERIMENTAL EVALUATION: 4 (sufficient evaluation or theoretical paper)

Comments of Reviewer 4

SUMMARY AND CONTRIBUTIONS: The paper describes a known weakness of LSB embedding, and proposes two counter measures.

OVERALL EVALUATION:
4 (borderline)

COMMENTS ON OVERALL EVALUATION:
The main idea is simple and interesting. However, there are many methods proposed in the past few years and I'm not sure whether the method described here have been studied before. Furthermore, there is a problem with the boundary cases, pixels with value 0 and 255. Using the proposed method will create artifacts that look like salt and pepper noise.

ORIGINALITY: 2 (minor originality)
REFERENCE TO PRIOR WORK: 3 (refer)
RELEVANCE/IMPORTANCE TO PSIVT: 3 (of sufficient interest)
CLARITY OF PRESENTATION: 3 (is clear enough)
TECHNICAL CORRECTNESS: 3 (probably correct)
EXPERIMENTAL EVALUATION: 4 (sufficient evaluation or theoretical paper)

PSIVT 2009 received 247 submissions, and accepted 40 papers for oral presentations and 58 for poster presentations. The acceptance rate is slightly less than 40%.
 

Monday, October 06, 2008

Message from PSIVT 2009

10月3日終於接到來自 PSIVT 2009 的消息, 接受了我們投稿的論文。可以開始著手準備前往東京了。

Dear Yeuan-Kuen Lee,

We wish to congratulate you on the acceptance of your submission with paper ID: 154 as an oral presentation in the PSIVT2009 program. Review comments for your paper are now available in the papers management system: https://cmt.research.microsoft.com/PSIVT2009.

The camera-ready paper deadline is on 17 October 2008, and paper preparation instructions can be found at http://psivt2009.nii.ac.jp/node/39. Please note that for your paper to be published in this conference, the camera-ready paper must be received by the deadline, and one of the authors must register for the conference by this deadline as well.

We are looking forward to the presentation of your paper in PSIVT2009.

Best regards,
PSIVT2009 Program Committee

Wednesday, September 24, 2008

Bivariate Distribution & Marginal Distribution

Bivariate Distribution 顧名思義就是具有兩個隨機變數的分配。

From StegoRN

貓頭鷹出版社 所出版的 統計學辭典 中, 舉的例子非常容易理解: 特定構造與型示的二手車。一般來說, 對有興趣的買主來說, 二手汽車有兩個令人感興趣且容易測量的變數: 車齡 里程數

假設有一家二手車行購進了同一樣式的二手車 30 輛, 車行就可以將其依車子的使用年齡和行駛里程數製成一個二變量次數表。換句話說, 就是列出一個二維表格, 一維是車齡, 另一維是里程數, 表中統計符合條件的車輛數。例如: 使用 2-3 年, 行駛 4-5 萬英里的車子有 2 輛。在這個例子中, 兩個變數並不是相互獨立的, 較舊的車通常行駛較長的距離。

From StegoRN

如果表中所使用的是機率, 那就是一個 二變量機率分布 (bivariate probability distribution), 在二變量分配表上, 分別做行相加, 或列相加的動作, 所得出的分配就稱為 邊際分配 (marginal distribution)

在二手車的例子中, 邊際分配為:
里程數  0-10 10-20 20-30 30-40 40-50 50-60 60-70 70-80
車輛數  1   2   2   3   7   6   6   3

車 齡  0 - 1   1 - 2   2 - 3   3 - 4   4 - 5   5 - 6   6 - 7   7 - 8
車輛數  3   2   7   9   3   3   1   2

維基百科中的 聯合分布 ( joint probability distribution ) 條目的說明, 其實和這邊的 Bivariate Distribution 的解釋是差不多的, 指的應該就是同樣的東西。

Tuesday, September 23, 2008

Feature-Based Steganalysis for JPEG Images and Its Implications for Future Design of Steganographic Schemes


Author: Jessica Fridrich

Information Hiding Workshop 2004
Toronto, Ontario, Canada
23 - 25, May, 2004

Lecture Notes in Computer Science, Vol. 3200



Abstract

In this paper, we introduce a new feature-based steganalytic method for JPEG images and use it as a benchmark for comparing JPEG steganographic algorithms and evaluating their embedding mechanisms. The detection method is a linear classifier trained on feature vectors corresponding to cover and stego images. In contrast to previous blind approaches, the features are calculated as an L1 norm of the difference between a specific macroscopic functional calculated from the stego image and the same functional obtained from a decompressed, cropped, and recompressed stego image. The functionals are built from marginal and joint statistics of DCT coefficients. Because the features are calculated directly from DCT coefficients, conclusions can be drawn about the impact of embedding modifications on detectability. Three different steganographic paradigms are tested and compared. Experimental results reveal new facts about current steganographic methods for JPEGs and new design principles for more secure JPEG steganography.

Monday, September 08, 2008

WCE 2008

2008/07/02

World Congress on Engineering (WCE) 每年都在 Imperial College, London 舉辦, 底下聯合了 15 個不同領域的研討會, 今年我們一行人投稿了兩篇論文, 一篇屬於 ICISIE'08, 由元智資工系王任瓚教授在 07/02 早上場次報告; 另一篇則是屬於 ICSIE'08, 由黃世育教授在 07/02 下午場次報告。


Imperial College, London
 

Imperial College, London

很幸運地, 出了地鐵站, 就遇見了已經在 WCE 2008 奮鬥了一個早上的三位教授。王任瓚教授已經報告完了, 所以一臉輕鬆。進入會場後, 我們先找個合適的地方留影,  


(left to right) Prof. Ran-Zan Wang, Prof. Shih-Yu Huang, Prof. Graeme Bell and I.
WCE 2008, Imperial College, London


Prof. Graeme Bell and I.
WCE 2008, Imperial College, London


Prof. Ran-Zan Wang and I.
WCE 2008, Imperial College, London


WCE 2008, Imperial College, London

最辛苦的兩位教授, 王任瓚教授已經報告完了, 所以一臉輕鬆。


Prof. Ran-Zan Wang and Prof. Shih-Yu Huang,
WCE 2008, Imperial College, London

會場提供的杯子是 IKEA 的 ... 便宜, 環保, 又非常好看!


Prof. Graeme Bell and Prof. Ran-Zan Wang
WCE 2008, Imperial College, London

這位學者看到我拿 Nikon D70 相機, 就跑來和我們閒聊 ...


WCE 2008, Imperial College, London

黃世育教授上台報告, ...


Prof. Shih-Yu Huang, WCE 2008, Imperial College, London


WCE 2008, Imperial College, London

看到窗外的大樹, 突然覺得可以在這邊上課是很幸福的 ...


Big trees outside the windows,
WCE 2008, Imperial College, London


贊助書商擺的攤子...


WCE 2008, Imperial College, London


WCE 2008, Imperial College, London


WCE 2008, Imperial College, London
 

Sunday, August 31, 2008

An Implementation of Key-Based Digital Signal Steganography

Author: Toby Sharp

Information Hiding Workshop 2001
Pittsburgh, PA, USA, April 25–27, 2001

Spring LNCS 2137, pp. 13-26, 2001

Abstract
A real-life requirement motivated this case study of secure covert communication. An independently researched process is described in detail with an emphasis on implementation issues regarding digital images. A scheme using stego keys to create pseudorandom sample sequences is developed. Issues relating to using digital signals for steganography are explored. The terms modified remainder and unmodified remainder are defined. Possible attacks are considered in detail from passive wardens and methods of defeating such attacks are suggested. Software implementing the new ideas is introduced, which has been successfully developed, deployed and used for several years without detection.

在談論這篇論文所使用的嵌入技術之前, 必須要先弄清楚 Pseudo-Random Sequence Generator 的運作方式。作者運用一個 Linear Feedback Shift Register (LFSR) 及使用傳訊者(sender) 和接收者(receiver) 所共同擁有的 stego-key 來建立與初始化這個 LFSR, 產生一連串的擬隨機序列。

使用 LFSR 所產生的擬隨機序列, 就可以決定要將訊息嵌入到 cover signal 的哪些 samples 之中, 作者稱造訪次序(visiting order) 為 sample sequence。假設 cover signal 一共有 P 個 samples, 令 t = log2(P), 因此我們就可以用 LFSR 所產生的 t 個位元來代表一個數字 k , 所以, 下一個要造訪的就是第 k 個 sample。作者在論文中還提到: 如果已經嵌入 x 位元, 剩下 (P-x) samples, 因此只要使用 t' = log2(P-x) 個擬隨機位元就可以決定下一個要造訪的 sample 了。

由上面的討論我們知道 sample sequence 是和 stego-key 值相關的, stego-key 不同, 所造訪的次序就不相同。除了和 stego-key 相關, 為了讓 sample sequence 也與 cover signal 及 embedded data相關, 作者還將所造訪的 sample value 的 Most significant Bit (MSB) 及 Least Significant Bit (LSB) 分別取出, 做為下一個擬隨機次序的前兩個位元。 而 MSB 所代表的就是 sample value, LSB 所代表的就是 embedded data。

這篇論文最重要的核心嵌入技術描述在 p. 15 的最後一段:
When a sample is visited, its data value is modified so that its least significant bit (LSB) is equal to the next bit of the secret data. The LSBs are not simply replaced; instead the whole sample value is incremented or decremented if the LSBs differ. This avoids the "pairs of values" statistical attack introduced in [9]. At each sample, one operation bit is taken from the generator and, if required, is used to determine whether to increment or decrement the sample value.
很可惜地, 論文中提到的隱藏工具 Hide, 網路上已經找不到下載點了。A. WestfeldIHW 2002 年所提出的論文 "Detecting Low Embedding Rates" 中有一張執行的初始畫面(Fig. 6),

 
在本篇論文的 Fig.4 則是展示使用者介面。 
 
 

Monday, August 11, 2008

Steganography in Cuil

From StegoRN

今天在 天下知識網Podcasting 聽到一篇文章 "為什麼全世界媒體都在報導這個網站?" 原來在介紹一個新的搜尋引擎 Cuil,...



我特地搜尋了 steganography 這個關鍵字, 搜尋結果的頁面排版還包含相關圖片, 感覺還真的不錯。仔細看搜尋結果, 包含的範圍似乎還蠻廣的, 各代表類型都有, 大家可以去試試看。

Sunday, August 03, 2008

LaTeX

由於 PSIVT 2009 並不鼓勵使用 MS Word 排版論文投稿, 因此我們開始準備將論文用 LaTeX 來編輯。雖然, 自己還在當博士生的時代, 曾經用 LaTex 寫過論文, 不過至少也是 6 年多前的陳年舊事了。以前用的系統, 早就老舊不堪, 也換了電腦。因此, 兩天前我開始在網路上搜尋, 準備重新建立 LaTex 的論文編輯環境。

最早搜尋到的一篇正體中文文章是我們互動媒體實驗室的新成員 小葉老師 在博士生時代 (2003) 所寫的 LaTeX 快速入門教學, 小葉老師推薦 WinShell + MiKTeX, 因此, 問題解決了一半, 直接上網搜尋這兩個套裝軟體的相關訊息, 開始研究。

以下是一些會使用到的相關軟體的網站:

MiKTeX is an up-to-date TeX implementation for the Windows operating system...

WinShell is a free multilingual integrated development environment (IDE) for LaTeX and TeX...

Ghostscript, Ghostview and GSview

CTAN: the Comprehensive TeX Archive Network



網路上還有一些寫得很不錯的中文文章, 值得推薦給大家:

1. 大家來學 LaTeX , by 李果正 Edward G.J. Lee

2. tw.bbs.comp.tex FAQ

PSIVT 2009: The 3rd Pacific-Rim Symposium on Image and Video Technology 2009

The 3rd Pacific-Rim Symposium on Image and Video Technology 2009
Tokyo, Japan, January 13th - 16th, 2009
Official Website: http://psivt2009.nii.ac.jp/

PSIVT 2009 is the continuation of a series of successful events in Hsinchu, Taiwan in 2006 and Santiago, Chile in 2007.
The symposium provides a forum for presenting and exploring the newest research and development in image and video technology by discussing the possibilities and directions in this field, and a place where both academic research and industrial activities are presented and meet for mutual benefit.

Main Themes
1. Image Sensors and Multimedia Hardware
2. Graphics and Visualization
3. Image and Video Analysis
4. Recognition and Retrieval
5. Multi-view Imaging and Processing
6. Computer Vision Applications
7. Video Communications and Networking
8. Multimedia Processing
 Image and video watermarking, steganalysis, steganography,
 multimedia content analysis, multimedia feature extraction, etc.

Important Dates:
Submission Deadline: August 11, 2008
Notification: September 22, 2008
Camera-ready Submission: beginning of October, 2008
Symposium: January 13–16, 2009

Call for papers

Sunday, March 30, 2008

Defending Against Statistical Steganalysis (part 1)

N. Provos10th USENIX Security Symposium, August 13-17, 2001 發表了 "Defending Against Statistical Steganalysis" 這篇論文, 內容就是闡述 OutGuess 0.2 這個隱藏軟體是如何運作的。

本篇文章所要討論的主軸是論文中有關 OutGuess 核心技術的部分 - Section 3 。

Section 3 Embedding Process

作者將 embedding Process 切割成兩個獨立的步驟:

1. Identification of redundant bits.
Redundant bits can be modified without detectably degrading the cover medium.
作者指出所謂的冗餘位元(redundant bits) 就是經過修改也不會在掩護媒體中產生會被偵測出來的品質下降現象(degrading)。

2. The selection of bits
in which the hidden information should be placed.

切割成兩個步驟的好處是容易取代(easy replacement), 如果要將本篇論文提出的方法在別的資料格式中實作出來, 只要將 identification algorithm 換掉, 然後用新的選擇策略(selection strategy)即可。

Section 3.1 Identification of Redundant Bits

作者闡述了一個觀念, 用來嵌入機密訊息的冗餘位元通常和影像的儲存格式相關。整個嵌入程序自然也和輸出格式有關。通常壓縮程序也包含其中。要最小化對掩護媒體(cover-medium)的修改(modification), 必須具備有關冗餘位元的相關知識才做得到, 作者提到 OutGuess 實作了整個輸出影像的運算。
For example, the OutGuess system performs all operations involved in created the output object and saves the redundant bits encountered. For the JPEG image format, this might be the LSB of the discrete cosine transform coefficients.

Section 3.2 Selection of Bits

探討如何從影像的 redundant bits 中選取一些 bits 來嵌入機密訊息。OutGuess 是使用 RC4 串流加密器(stream cipher)對機密訊息加密, 同時也用 RC4 來建立一個 PRNG (pseudo-random number generator), 然後再將選定的 seed 餵進這個 PRNG 來選擇冗餘位元。

32 state bits = 16-bit seed + 16 bit integer
16-bit seed: 由於不同的 seeds 會選取不同的冗餘位元來作為嵌入機密訊息之用, 因此, 不同的 seeds 自然對原始影像造成的 change, 也會有所不同。當接收端(receiver)收到偽裝影像(stego-image)後, 必須知道當初所選定的 seed, 因此必須把這16-bit seed 也嵌入到掩護影像(cover-image) 之中。
16-bit integer: containing the length of the hidden message.

冗餘位元的選取方式是利用上述的 PRNG 來計算下一個 bit 的隨機距離(random offset) R i(x),

 b0 = 0,
 bi = bi-1 + Ri(x)  for i = 1, 2, ... , n

bi 表示第 i 個選取位元的位置, Ri(x) 表示與上個選取位元之間的隨機距離, 值介於 [1, x] 之間。x 為最大的間隔(interval), 這個值在每嵌入 8 個位元, 就會重新使用下列的公式重新計算, 目的就是讓所有的機密訊息可以分布到整個可以使用的位元中。

 interval = 2 * remaining redundant bits / remaining length of message.

用上述的方法來設定 interval, 會使得機密訊息的長度限制在 50% 嵌入空間之內。

Section 3.3 Beneficial Reseeding of the PRNG


談論如何靠著選擇不同的 seeds, 智慧地選擇不同的嵌入位置的子集合, 不但可以讓 changed bits 的總數降低, 而且使得嵌入行為較不容易被偵測出來 (Detectability is also used as a bios in the selection process.)。

由於掩護影像(cover-image)中的冗餘位元, 不是 1 就是 0, 加上要嵌入的資料先用 RC4 stream cipher 加密, 變成一串二元的隨機資料流(binary random stream), 將機密訊息嵌入到冗餘位元, 造成這些冗餘位元被改變的機率期望值為 0.5。因此, 統計學中的二元分布(binomial distribution)正好可以用來描述一般的 LSB 嵌入行為。

假設, 我們從冗餘位元之中, 將一個 seed 餵進 PRNG 選擇了 4430 個位元, 並將同樣長度的機密訊息嵌入其中, 便可以去計算此次嵌入動作一共改變了多少個 redundant bits。注意: 不同的 seed 餵進同一個 PRNG 將使得所選擇的嵌入位置不同。Figure 1 就是重複使用不同的 seeds 來統計這 4430 個redundant bits 被改變的總數, 累計其統計值所畫出來的結果。


Figure 1: Probability distribution of changed bits for different seeds compared to a binomial distribution with n=4430 and p=0.5.

不管是從 binomial distribution 公式推論, 或是從 Figure 1 中的實驗中, 我們都可以觀察到當我們選定一個 seed 時, changed bits 的個數是以 n/2 = 2215 的可能性(機率)最高, 不過, 還是存在一些 seeds 會使得 changed bits 的個數小於 2150。論文中是這樣討論的:
Picking a seed that represents the changed bits at the lower end of the binomial distribution allows us to reduce the number of bits that have to be changed; see Figure 1. It becomes harder to detect the modifications, as more of the hidden message is already naturally represented in the redundant bits.
除了降低修改之外, 可偵測性(detectability)也是 selection process 要考量的一個因素。
Detectability is also used as a bias in the selection process. The selector does not try to reduce only the number of changed bits but also the overall detectability. Whenever a bit has to be modified, its detectability will be added to a global bias. A higher accumulated bias reduces the likelihood that this specific embedding will be used.

Section 3.4 Choices with Coding Theory

作者在這邊提到 Coding Theory 的考量為使用 PRNG 去選擇冗餘位元就無可避免地選到
1. locked bits
2. bits with a high detectability
上述兩類冗餘位元是作者不想去更改的。因此, 作者想使用錯誤更正碼(error-correcting codes)來解決上述問題。

[n, k, d] coding 指的是長度為 k 位元的機密訊息(k-bit data block), 將被編碼成長度為 n 位元的編碼區塊(n-bit code block), 每個 code 之間的 Hamming distance 至少是 d, 假設 d = 2t + 1, 那麼這個編碼就具備了可以更正 t 個錯誤位元的能力。換句話說, n 個位元的編碼區塊之中, 如果發生 t 個位元的錯誤, 那麼使用解碼程序, 就可以偵測出哪 t 個位元發生錯誤, 進而更正回來, 因此原先的 k 位元的資料, 是可以完全解碼出來的。

將機密訊息用錯誤更正碼來編碼, 無疑也會增加要嵌入的長度。然而, 觀察整個嵌入過程獲知:
1. 有一半的資訊嵌入是不會改變到冗餘位元的 ( n / 2),
2. 可以有 t 個位元可以不用嵌入(更改冗餘位元)
因此, 假如
 ( n /2 ) - t = ( k / 2 )
成立, 那麼作者希望嵌入 n 位元的編碼區塊需要修改的位元數(上述式子的等號左邊)要和嵌入未經編碼的 k 位元的資料區塊需要更改的位元數(上述式子的等號右邊)相等。將上述式子通分得到 n -2t = k, 並將 d+1 = 2t 帶入可以得到
 d = n - k + 1,
剛好就是 MDS (maximum distance separable) code 的 Singleton bound。因此, 作者在這邊得到一個結論就是只要選擇 MDS codes, 就可以滿足上述作者期望的。

不幸地, 值得一提的(non-trivial)二元 MDS code 就只有重複碼(repetition code), 主要缺點就是編碼必須將資料重複 n 次, 因此, 重複碼僅使用在隱藏訊息很小的情況。

資料經過編碼後, 每個編碼區塊可以選擇 t 個位元不去修改冗餘位元。作者所使用的選擇策略是:
欲嵌入的位元與冗餘位元不同(conflict), 且冗餘位元先前已經被嵌入過資料, 被 locked bits 鎖住。

OutGuess 所使用的技術與 Ross J. Anderson and Fabien A. P. Petitcolas 發表在 Journal on Selected Areas in Communication, 16(4): 474-481, May, 1998 的論文 On the Limits of Steganography 中所建議的 parity encoding 相似。然而, 使用 error-correcting codes 的好處要比使用 parity encoding 多。透過選擇一種不是 MDS 的 code, 我們可以犧牲些許的嵌入容量(capacity) 而得到更高的安全性(security)。除此, 對照使用 parity encoding 必須 lock 住 n 個位元, 使用 error-correcting codes 則僅僅需要 lock 住 n-t 個位元。

Section 3.5 Plausible Deniability

為了嵌入機密訊息, 我們修改掩護媒體中的冗餘位元。這些冗餘位元可能存在一些我們沒有感知, 或是對手比我們了解的自然統計特質。假如嵌入程序改變了上述特質, 在這方面知識淵博的觀察者, 不用指出哪些特定位元被改變, 就可以推論出隱藏訊息是存在的。

偽裝媒體的創造者必須面對的是: 欲隱藏的通訊行為可能被揭露出來。然而, 我們假設觀察者僅僅可以確定的事實是掩護媒體被更改了。假如傳訊者嵌入多重訊息, 其中可以包含一份無害的訊息, 讓它和真正想要傳送的訊息(request)攪在一起, 然後宣稱沒有任何訊息隱藏在偽裝媒體之中, 偽裝媒體並沒有遭受破壞(沒有遭到修改, 換句話說就像原始掩護影像一樣, 沒有破壞原先存在的特質)。這就是所謂的似乎合理的可否認性(plausible deniability)

  

實際上, 整個 Section 3 所描述的技術已經隱含地支援上述所提到的似乎合理的可否認性。可以隱藏不只一份的訊息, 使用 locked bits 來避免先嵌入的訊息被後嵌入的訊息覆蓋掉。即使是與嵌入訊息的大小相關, 不與先前 locked 住的冗餘位元重疊的可能性是很小的, 在這種情況下, 使用 error-correcting codes 則是可以增加選擇的彈性。

Section 3.6 Hidden Message Determines Cover

針對特定的隱藏訊息, 可以在不同的掩護媒體中, 選擇一個機密訊息對掩護媒體本身影響較小(with minimal modification)的掩護媒體, 來嵌入機密訊息。這和 Section 3.3 中有關 binomial distribution 的系列討論是差不多的。

Go to: Defending Against Statistical Steganalysis (part 2)

Go to: Defending Against Statistical Steganalysis (part 3)

Niels Provos, "Defending Against Statistical Steganalysis,"10th USENIX Security Symposium, August 13-17, 2001.
 

Wednesday, March 19, 2008

The Difference between Conference and Workshop

昨天上課在討論時, 研究生問我 conference 和 workshop 有什麼不同? 我記得前幾個禮拜曾經在空中英語教室 chatroom 聽過這麼一段討論, 於是就把檔案找出來一起聽聽看。剛剛想到和黃世育老師規劃非資訊學院科系大二多媒體課程時, 有納入音訊檔的簡單處理, 黃老師建議使用 GoldWave , 於是就上網搜尋這個軟體下載, 想要把 chatroom 這一段單獨剪輯成一個檔案, 放在這邊給研究生分享。


Sunday, March 16, 2008

IHW 2008: Information Hiding Workshop 2008

Information Hiding Workshop 2008
Santa Barbara, California, USA,
May 19-21, 2008

For many years, information hiding has captured the imagination of researchers. Digital watermarking and steganography tools are used to address digital rights management, protect information, and conceal secrets. From an investigators perspective, information hiding provides an interesting challenge for digital forensic investigations and steganalysis techniques allows hidden information to be discovered. These are but a small number of related topics and issues. Current research themes include:

  * Watermarking (attacks, security, algorithms)
  * Anonymity and privacy
  * Steganography and steganalysis
  * Multimedia security
  * Other hiding domains (e.g. text, software, etc.)
  * Information assurance
  * Digital forensics
  * Covert/subliminal channels
  * Theoretical aspects of information hiding
  * Intrusion detection
  * Digital rights management
  * Novel technologies/applications

Continuing a successful series that brought together these closely linked research areas, the 10th Edition of Information Hiding (IH08) will be held in Santa Barbara, California.
Call for papers

Saturday, March 15, 2008

EDAS Conference Management System

要投稿到 ISC 2008, 必須透過 EDAS Conference Management System, 換句話說, 我們必須要先到 EDAS 註冊, 取得帳號後, 才能投稿論文。這兩天, 使用 EDAS 的感覺還不錯, 投稿時, 要求論文初稿上不能有作者的相關資料, 以維護審稿的公正性, 因此每一個作者也都必須要有 EDAS 帳號, 然後再用新增作者的功能加上去。

有許多研討會都是透過 EDAS 來投稿論文, 我上星期剛接觸到 EDAS 時, 在 EDAS 的網站看到許多可以投稿的研討會相關資訊, 那時候我就覺得透過這個系統中 Submit paper 功能, 看到原本分散在各處的研討會訊息被整合起來, 對每年都要參加研討會的學者來說, 應該是很不錯的管道。

Tuesday, March 11, 2008

Message from ISC 2008

今天中午收到來自 ISC 2008 的 mail, 說明目前投稿狀況, 台灣只有 8 篇, 真的是有點少, 希望大家多多投稿到 ISC 2008,...

收到的 mail 如下:

各位好,
TWISC 與資訊安全學會將於九月主辦第十一屆 Information Security Conference 國際會議,李德財所長擔任會議之 General Chair,論文集列入 LNCS,目前投稿情形有五十六篇,其中就日本投了超過十七篇,我們主辦國卻只有八篇投稿,截稿日期到 3/15 日截止,會議網站請參考:http://isc08.twisc.org/ 若各位目前手頭有稿件,祈請踴躍投稿。

PS:依學術慣例,LNCS 接受之論文,若有進行 Major revision,也可再轉投 journal。倘若論文(初稿)沒有被接受,Reviewer 的意見,也將對未來論文修改有許多助益。

吳宗成 敬上

Thursday, March 06, 2008

ISC08: Information Security Conference 2008

September 15-18, 2008
Taipei, Taiwan
Website http://isc08.twisc.org/index.html

Information Security 每年一度的盛事, 今年剛好在台灣舉辦, 研討會的 Topics 包含了我們的研究領域 - Information Hiding, ( 另一個重要的國際研討會是 IHW, Information Hiding Workshop, 今年 May 19-21 在 Santa Barbara, California, USA 舉辦, 論文截稿日期是 February 2, 2008 早就已經 來不及了 )。

今年 ISC 原本的 Submission Deadline 是 March 1, 2008, 我的論文原先的規劃是準備投稿到屬於 WCE 2008ICISOIE 08 , 昨天其實已經將論文改成 WCE 2008 的六頁格式, 準備今天投稿出去。早上開車到學校的路上, 就一直在考慮改投到 ISC 08, 在車上和黃世育老師討論後, 決定將我們一起合作寫的兩篇論文分散到兩個研討會, 黃老師寫的 CAPTCHA 那篇, 照原定計畫投稿的 WCE 2008, 我寫的這篇 Modified LSB Embedding 改投到 ISC 08, 兩個研討會都要參加。

昨天花了一整天, 將論文調整成 WCE 2008 的格式, 決定改投之後, 今天又要把論文格式重新改成 ISC 08 的格式了。

Call for papers


Important Dates
Conference Dates: September 16-18, 2008
Submission Deadline: March 15, 2008, 11pm GMT (firm deadline)
Notification of Acceptance: May 20, 2008
Camera-ready Copies Deadline: June 15, 2008

Topics of Interest
ISC aims to attract high quality papers in all technical aspects of information security.
Topics of interest include, but are not limited to, the following:

 * Access Control
 * Accounting and Audit
 * Anonymity and Pseudonymity
 * Applied Cryptography
 * Attacks and Prevention of Online Fraud
 * Authentication and Non-repudiation
 * Biometrics
 * Cryptographic Protocols and Functions
 * Database and System Security
 * Design and Analysis of Cryptographic Algorithms
 * Digital Rights Management
 * Economics of Security and Privacy
 * Formal Methods in Security
 * Foundations of Computer Security
 * Identity and Trust Management
 * Information Hiding and Watermarking
 * Infrastructure Security
 * Intrusion Detection, Tolerance and Prevention
 * Mobile, Ad Hoc and Sensor Network Security
 * Network and Wireless Network Security
 * Peer-to-Peer Network Security
 * PKI and PMI
 * Private Searches
 * Security and Privacy in Pervasive/Ubiquitous Computing
 * Security in Information Flow
 * Security for Mobile Code
 * Security of Grid Computing
 * Security of eCommerce, eBusiness and eGovernment
 * Security Modeling and Architectures
 * Security Models for Ambient Intelligence environments
 * Trusted Computing
 * Usable Security
 

Saturday, November 24, 2007

Structure of AC code table

在 JPEG 規格書 P. 89 談到了 AC code table 的結構。每一個非零的 AC 係數都是以一個 8-bit RS (Run/Size) 的形式來描述。
RS = binary 'RRRRSSSS'
4 個低位元 SSSS 定義非零的 AC 係數所屬的類別 (category), 分成 10 個類別, 可由 Table F.2 查到每個類別的範圍。高位元的 4 個 RRRR 則是指出非零 AC 係數之前存在多少個 0。由於有可能超過 15 個 0, 4-bit 的 RRRR 無法表示, 因此定義了 'RRRRSSSS' = X'F0' 來代表 15 個 0 外加一個值為 0 的 AC 係數, 換句話說, 共16個 0。除此, 如果在 Zigzag 序列中, 後面的 AC 係數已經全部為 0 了, 就用一個 EOB (end-of-block), 'RRRRSSSS'= '00000000' 來表示。

Table F.2 Categories assigned to coefficient values
SSSS   AC coefficients
1        -1, 1
2        -3,-2, 2, 3
3        -7..-4, 4..7
4       -15..-8, 8..15
5       -31..-16, 16..31
6       -63..-32, 32..63
7      -127..-64, 64..127
8      -255..-128, 128..255
9      -511..-256, 256..511
10     -1023..-512, 512..1023

Huffman encoding procedures for DC coefficients

在 JPEG 影像壓縮標準中, 對於 DC 係數的編碼, 霍夫曼編碼程序(Huffman encoding procedures) 使用了下列 2 個擴展表格(extended tables):
  1) XHUFCO
  2) XHUFSI
這兩個表格都是以相鄰兩個 block 區塊的差值(DIFF) 為索引, XHUFCO 存放 DIFF 所屬類別(category) 的 Huffman code, 再接上 為了區分 DIFF 在此類別中的位置的附加位元(additional bits)。XHUFSI 則是存放 DIFF 值所對應的 XHUFCO 表格中的位元長度, 即 Huffman code 長度加上附加位元的長度。

XHUFCO 與 XHUFSI 這兩個表格則是由 Annex C, P. 52 所談到的 EHUFCO 與 EHUFSI 兩個表格加上附加位元擴展而來。

有了 XHUFCO 與 XHUFSI 兩個表格, 當我們有一個 DIFF 需要編碼時, 只要快速查表即可很快地產生 binary data。針對 DC 差值 (DIFF), 霍夫曼編碼程序如下:

  SIZE = XHUFSI(DIFF)
  CODE = XHUFCO(DIFF)
  code SIZE bits of CODE
 

Wednesday, November 14, 2007

Huffman decoding of DC coefficients

由於以下兩個原因:

1) DCT 的 DC 係數就是代表 8*8 區塊的像素平均值,
2) 影像中相鄰區塊的像素平均值有很大的機率是相近的

因此 JPEG 的 Huffman encoding procedure 並不是直接對 DC coefficients 編碼, 而是對兩個相鄰區塊的 DC difference 編碼。由於 difference 可能的值涵蓋太廣了, 對每一個可能值採取直接編碼的方式並不可行 (decoding tree 太大了), 因此 JPEG 採用的作法是先對所有可能值分成12個類別(category), 然後統計每一類別的機率, 有了機率後, 再依機率做 Huffman 編碼。JPEG 規格書 P.89 的 Table F.1 列出了 DC difference 是如何分類的

Table F.1 Difference magnitude categories for DC coding
SSSS   DIFF values
0          0
1        -1, 1
2        -3,-2, 2, 3
3        -7..-4, 4..7
4       -15..-8, 8..15
5       -31..-16, 16..31
6       -63..-32, 32..63
7      -127..-64, 64..127
8      -255..-128, 128..255
9      -511..-256, 256..511
10     -1023..-512, 512..1023
11     -2047..-1024, 1024..2047

除了第 0 類 (SSSS=0) 之外, 每一類中, difference 的可能值並不只有一個, 因此還另外需要一些額外的位元(appended additional bits) 來標示出確切的 difference 值。所需要的額外位元總數就剛好是 SSSS 值, 原因很簡單, 第 1 類 (SSSS=1) 中有 -1, 1 兩個可能值, 因此只需要 1 個額外位元來指出真正的 difference 值是哪一個。同理, 第 2 類 (SSSS=2) 中有 -3, -2, 2, 3 共 4 個可能值, 因此需要 2 個額外位元來指出真正的 difference 值是哪一個。其餘類別的個數剛好就是 2^SSSS 個, 因此共需要 SSSS 個額外位元才夠指出真正的 difference 值為何。

如果 DIFF > 0, 額外位元就是 DIFF 的二進位表示法中的 SSSS 個低位元(low order bits), 如果 DIFF < 0, 則是用 (DIFF-1) 的 SSSS 個低位元。 我們舉例說明: a) 假設 DIFF 值為 -9, SSSS=4, 因為DIFF < 0, 因此 (DIFF-1) = -10, 用 2-complement 來表示 -10, 為 11110110, 那麼 4 個低位元就是 0110。 JPEG 解壓縮時, Huffman decoding procedure 如果使用 decoding tree 解碼得到 SSSS=4, 就會繼續取得 4 個額外位元 0110, 根據第 1 個位元判斷 DIFF<0, 然後使用後面 3 個位元計算得到 6, SSSS=4 這個類別負數區間的最小值為 -15, 將 -15 + 6 = -9 , 就可以得到 DIFF = -9。 b) 假設 DIFF 值為 6, SSSS=3, 因為 DIFF > 0, 因此用 2-complement 來表示 6 為 00000110, 那麼 3 個低位元就是 110。

使用 decoding tree 得到 SSSS=3 後, 繼續取得 3 個位元 110, 由於第 1 個位元為 1, 因此判斷 DIFF 屬於正數區間, 後面兩個位元為 10, 值為 2, 正數區間的最小值為 4, 4 + 2 =6 因此得到 DIFF=6。 

Saturday, November 10, 2007

Y'CbCr in JPEG Standard

JPEG allows Y'CbCr where Y', Cb and Cr have the full 256 values:

 JPEG-Y'CbCr (601) from "digital 8-bit R'G'B' "
 ===========================================
 Y' = + 0.299 * R'd + 0.587 * G'd + 0.114 * B'd
 Cb = 128 - 0.168736 * R'd - 0.331264 * G'd + 0.5 * B'd
 Cr = 128 + 0.5 * R'd - 0.418688 * G'd - 0.081312 * B'd
 ===========================================
 R'd, G'd, B'd in {0, 1, 2, ..., 255}
 Y', Cb, Cr in {0, 1, 2, ..., 255}

from Wikipedia: YCbCr

Monday, November 05, 2007

關於 JPHide 的點點滴滴 (一) : N. Provos

Niels Provos 在 "Detecting Steganographic Content on the Internet" 這篇論文的 Section 5.2 整節都在談論 JPHide 這個隱藏軟體。內文如下:
 JPHide is a steganographic system by Allan Latham. There are two versions: 0.3 and 0.5. Version 0.5 supports additional compression of the hidden message. As a result, they use slightly different headers to store embedding information. Before the content is embedded, it is Blowfish encrypted with a usersupplied pass phrase.

Because the DCT coefficients are not selected continuously from the beginning, JPHide is more difficult to detect.

The program uses a fixed table that defines classes of DCT coefficients to determine in which order to modify the coefficients. All coefficients in the current class are used first to hide information before the next class is chosen. As a result, coefficients are selected in such a way that they those likely to be numerically high are used first.

One artifact of the implementation is that the information hiding continues in the current coefficient class even after the complete message has been embedded. The first class in the table are the DC coefficients of color component zero. An image with a resolution of 600 * 480 has approximately five thousand DC coefficients. Even if the message is only eight bits long, JPHide modifies all five thousand coefficients in such an image.
這邊提到 JPHide 有一個特殊方式來定義嵌入次序, JPHide 使用一個固定的表格來將 DCT 係數分成不同的 classes, 整張影像相同 class 中的 DCT 係數會依序拿來嵌入機密訊息, 直到此 class 的係數用完了, 才會動用到下一個 class 的 DCT 係數。接著, 相同 class 的係數, 數值較大者也會優先拿來嵌入機密訊息。個人覺得這樣做是有道理的, 因為嵌入影響對較大值的係數來說, 比例相對較小, 因此優先使用。

另外一點令人匪宜所思的是: 即使所有的機密訊息已經嵌入完畢了, JPHide 依然會繼續修改目前這個 class 的所有 DCT 係數。論文中提到一個例子, 第一個 class 就是 DC 係數, 假設一張 600*480的影像, 就會有 (600/8)*(480/8)= 75*60 = 4500 個 DC 係數, 那麼即使機密訊息只有 8 bits, JPHide 依然會去修改這所有的 DCT 係數。
 A pseudo-random number generator determines if coefficients are skipped. The probability of skipping bits depends on the length of the hidden message and how many bits have been embedded already.

JPHide modifies not only the least-significant bits of the DCT coefficients, it can also switch to a mode where the second-least-significant bits are modified.
如其他軟體一般, JPHide 用一個 PRNG 來決定哪些係數該跳過不嵌入機密訊息。然而, 較特殊的作法是跳過的機率是和 1) 機密訊息的長度, 2) 已經嵌入多少資料量。這代表每嵌入 1 個位元, 機率值就隨時進行更新, 用以控制所有的訊息可以完全順利嵌入。另外, JPHide 也會將機密訊息嵌入到次低位元中。

From StegoRN
Figure 6: JPHide has a signature similar to JSteg. The major difference is the order in which the DCT coefficients are modified.
Figure 6 shows the probability of embedding for an image containing information hidden with JPHide. Because JPHide can skip DCT coefficients, the probability is not as high as with JSteg.
Figure 6 是使用 Chi-Square Attack 來針對 JPHide stego-images 分析, 橫軸是將影像平分成 100 等份, 每一等份都用 Chi-Square Attack 計算嵌入機率 p。由於 JPHide 會跳過部份的 DCT 係數不藏, 因此所得到的 P 值並不像 Jsteg 那麼高。


Niels Provos and Peter Honeyman, "Detecting Steganographic Content on the Internet,"ISOC NDSS'02, San Diego, CA, February 2002.